Ask Question to GPT
The final step of the RAG pipeline. Take the document chunks retrieved from Azure AI Search (via keyword, semantic, or vector search) and send them as context to the Azure OpenAI Chat Completions API. GPT generates a grounded answer based only on the provided context — not from its training data — making responses accurate and traceable to your enterprise documents.
| Property | Value |
|---|---|
| Method | POST |
| Content-Type | application/json |
| Service | Azure OpenAI (direct API call) |
| API Version | 2025-01-01-preview |
| Model | Your GPT deployment (e.g. gpt-4o, gpt-4). Set via {{openAI_GPTModelName}}. |
https://my-resource.openai.azure.com/gpt-4o
Azure OpenAI API Key
api-key{{openAI_Key}}| Field | Type | Required | Description |
|---|---|---|---|
| messages | Array | Required | Chat message array. Include a system message with the RAG instruction and a user message with the question + context chunks. |
| messages[].role | String | Required | "system" for the grounding instruction, "user" for the question + context. |
| messages[].content | String | Required | The message text. For the user message, format as: Question:\n{question}\n\nContext:\n{chunk1 content}\n{chunk2 content}... |
| temperature | Float | Optional | Randomness of the response. Use 1 (default) for balanced responses, 0 for deterministic answers. |
| max_completion_tokens | Integer | Optional | Maximum tokens in the response. Default in collection is 500. |
Request Body (from Postman collection)
{
"messages": [
{
"role" : "system",
"content": "You are an Enterprise Knowledge Assistant. Answer the user's question ONLY using the provided context. Do not use external knowledge. If the answer cannot be found in the context, reply exactly: 'I couldn't find that information in the provided documents.' Quote only when necessary and keep the answer concise."
},
{
"role" : "user",
"content": "Question:\nWhat is this document all about?\n\nContext:\n\n### Chunk 1: <section heading>\n<chunk content from search results>\n\n### Chunk 2: ...\n<chunk content>"
}
],
"temperature" : 1,
"max_completion_tokens" : 500
}content of the user message. Use the format ### Chunk {n}: {section}\n{content} for each chunk to give GPT clear structure.
cURL
curl -X POST \ -H "api-key: <your-azure-openai-key>" \ -H "Content-Type: application/json" \ -d '{"messages":[{"role":"system","content":"Answer ONLY using the context..."},{"role":"user","content":"Question:\nWhat is this about?\n\nContext:\n### Chunk 1: Intro\nAzure AI Search is..."}],"temperature":1,"max_completion_tokens":500}' \ "https://<resource>.openai.azure.com/openai/deployments/gpt-4o/chat/completions?api-version=2025-01-01-preview"
Sample Response
{
"id" : "chatcmpl-...",
"object" : "chat.completion",
"choices": [
{
"index" : 0,
"finish_reason": "stop",
"message": {
"role" : "assistant",
"content": "This document is the Norvia University Student Handbook for 2026-27. It covers academic regulations, fees, student services, and codes of conduct for all enrolled students."
}
}
],
"usage": {
"prompt_tokens" : 412,
"completion_tokens": 48,
"total_tokens" : 460
}
}choices[0].message.content.api-key.messages field, or context exceeds the model's context window limit.