← Back to APIs
RAG PIPELINE — STEP 4 (FINAL)

Ask Question to GPT

The final step of the RAG pipeline. Take the document chunks retrieved from Azure AI Search (via keyword, semantic, or vector search) and send them as context to the Azure OpenAI Chat Completions API. GPT generates a grounded answer based only on the provided context — not from its training data — making responses accurate and traceable to your enterprise documents.

POST · Azure OpenAI Chat Completions · GPT-4 / GPT-4o Variables: openAI_Endpoint, openAI_Key, openAI_GPTModelName
01

Endpoint

POST {{openAI_Endpoint}}openai/deployments/{{openAI_GPTModelName}}/chat/completions?api-version=2025-01-01-preview
PropertyValue
MethodPOST
Content-Typeapplication/json
ServiceAzure OpenAI (direct API call)
API Version2025-01-01-preview
ModelYour GPT deployment (e.g. gpt-4o, gpt-4). Set via {{openAI_GPTModelName}}.
02

Postman Collection Variables Required

Set these variables in the Postman collection before calling this API openAI_Endpoint — Your Azure OpenAI base URL, e.g. https://my-resource.openai.azure.com/
openAI_Key — Your Azure OpenAI API key
openAI_GPTModelName — Your GPT deployment name, e.g. gpt-4o
03

Authentication

Azure OpenAI API Key

Header: api-key{{openAI_Key}}
04

Request Body

FieldTypeRequiredDescription
messages Array Required Chat message array. Include a system message with the RAG instruction and a user message with the question + context chunks.
messages[].role String Required "system" for the grounding instruction, "user" for the question + context.
messages[].content String Required The message text. For the user message, format as: Question:\n{question}\n\nContext:\n{chunk1 content}\n{chunk2 content}...
temperature Float Optional Randomness of the response. Use 1 (default) for balanced responses, 0 for deterministic answers.
max_completion_tokens Integer Optional Maximum tokens in the response. Default in collection is 500.

Request Body (from Postman collection)

{
  "messages": [
    {
      "role"   : "system",
      "content": "You are an Enterprise Knowledge Assistant. Answer the user's question ONLY using the provided context. Do not use external knowledge. If the answer cannot be found in the context, reply exactly: 'I couldn't find that information in the provided documents.' Quote only when necessary and keep the answer concise."
    },
    {
      "role"   : "user",
      "content": "Question:\nWhat is this document all about?\n\nContext:\n\n### Chunk 1: <section heading>\n<chunk content from search results>\n\n### Chunk 2: ...\n<chunk content>"
    }
  ],
  "temperature"           : 1,
  "max_completion_tokens" : 500
}
How to build the Context Concatenate the search results from Steps 3, 4, or 6 into the content of the user message. Use the format ### Chunk {n}: {section}\n{content} for each chunk to give GPT clear structure.
05

Sample Request

cURL

curl -X POST \
  -H "api-key: <your-azure-openai-key>" \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"system","content":"Answer ONLY using the context..."},{"role":"user","content":"Question:\nWhat is this about?\n\nContext:\n### Chunk 1: Intro\nAzure AI Search is..."}],"temperature":1,"max_completion_tokens":500}' \
  "https://<resource>.openai.azure.com/openai/deployments/gpt-4o/chat/completions?api-version=2025-01-01-preview"
06

Response — 200 OK

Sample Response

{
  "id"     : "chatcmpl-...",
  "object" : "chat.completion",
  "choices": [
    {
      "index"        : 0,
      "finish_reason": "stop",
      "message": {
        "role"   : "assistant",
        "content": "This document is the Norvia University Student Handbook for 2026-27. It covers academic regulations, fees, student services, and codes of conduct for all enrolled students."
      }
    }
  ],
  "usage": {
    "prompt_tokens"    : 412,
    "completion_tokens": 48,
    "total_tokens"     : 460
  }
}
07

Error Responses

200
OK — GPT response returned in choices[0].message.content.
401
Unauthorized — Invalid or missing api-key.
404
Not Found — GPT deployment name not found in your Azure OpenAI resource.
400
Bad Request — Missing messages field, or context exceeds the model's context window limit.
429
Too Many Requests — Rate limit or token-per-minute quota exceeded for this deployment.