← Back to APIs
RAG PIPELINE — STEP 1

Chunk and Perform Embedding

Upload a PDF to the Lowcademy Labs endpoint. It extracts the text, splits it into 800–1200-character semantic chunks, optionally calls Azure OpenAI Embeddings to generate a vector per chunk, and returns a payload you can POST directly to Azure AI Search without any modification. This is the first step in the RAG document ingestion pipeline.

POST · Lowcademy Hosted Azure OpenAI Embeddings · PDF Processing Postman Variable: openAI_Endpoint, openAI_Key, openAI_EmbeddingModelName
01

Endpoint

POST https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php
PropertyValue
MethodPOST
Content-Typemultipart/form-data
ParametersAll parameters passed as request headers — not query params or form fields
File UploadPDF binary in the document multipart form field
ResponseAzure AI Search Documents–Index API payload: { "value": [...] }
02

Postman Collection Variables Required

Set these variables in the Postman collection before calling this API These variables are only needed when performEmbedding: true. If you skip embedding, no Azure OpenAI variables are required.

openAI_Endpoint — e.g. https://my-resource.openai.azure.com/
openAI_Key — Your Azure OpenAI API key
openAI_EmbeddingModelName — e.g. text-embedding-3-small
03

Authentication

Lab API Key — For Training Use Only

Header: api-keyankitg.in
HeaderTypeRequiredDescription
api-key String Required Fixed lab API key. Returns 401 if missing or incorrect.
04

Request Headers

HeaderTypeRequiredDescription
sourceName String Required Logical name for the document (e.g. AzureGuide). Written to every chunk's sourceName field.
tenant_id Int64 Required Tenant identifier — must be numeric. Stored in every chunk.
document_id Int64 Required Document identifier — must be numeric. Used to build the chunk id key.
performEmbedding Boolean Optional Pass true to generate Azure OpenAI embeddings per chunk. Defaults to false; embedding array will be [].
embeddingEndpoint String (URL) Conditional Your Azure OpenAI base URL. Required when performEmbedding: true. Maps to {{openAI_Endpoint}}.
embeddingKey String Conditional Azure OpenAI API key. Required when performEmbedding: true. Maps to {{openAI_Key}}.
embeddingDeploymentName String Conditional Embedding deployment name (e.g. text-embedding-3-small). Required when performEmbedding: true. Maps to {{openAI_EmbeddingModelName}}.
05

Request Body — File Upload

FieldTypeRequiredDescription
document Binary (PDF) Required The PDF file. Must contain selectable text (not a scanned/image-only PDF). Send as multipart/form-data.
06

Sample Request

cURL — with embedding (embedding array populated)

curl -X POST \
  -H "api-key: ankitg.in" \
  -H "sourceName: MyDocument" \
  -H "tenant_id: 1" \
  -H "document_id: 42" \
  -H "performEmbedding: true" \
  -H "embeddingEndpoint: https://my-resource.openai.azure.com" \
  -H "embeddingKey: <your-azure-openai-key>" \
  -H "embeddingDeploymentName: text-embedding-3-small" \
  -F "document=@/path/to/document.pdf" \
  https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php

cURL — without embedding (embedding will be [])

curl -X POST \
  -H "api-key: ankitg.in" \
  -H "sourceName: MyDocument" \
  -H "tenant_id: 1" \
  -H "document_id: 42" \
  -F "document=@/path/to/document.pdf" \
  https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php
07

Response — 200 OK

Ready for Azure AI Search The response body is the exact JSON you POST to Azure AI Search /docs/index. No wrapping or transformation needed. Copy and send directly.
FieldTypeDescription
valueArrayArray of chunk objects ready for Azure AI Search indexing.
value[].@search.actionStringAlways "upload".
value[].idStringUnique key: {tenant_id}-{document_id}-{zero-padded-index}, e.g. 1-42-0001.
value[].tenant_idInt64Echoed from the tenant_id header.
value[].document_idInt64Echoed from the document_id header.
value[].sourceNameStringEchoed from the sourceName header.
value[].sectionStringDetected heading for this chunk (max 90 chars).
value[].contentStringChunk text, 800–1200 characters.
value[].embeddingArray<float>1536-dim float array when performEmbedding: true. Empty [] otherwise.

Sample Response

{
  "value": [
    {
      "@search.action" : "upload",
      "id"             : "1-42-0001",
      "tenant_id"      : 1,
      "document_id"    : 42,
      "sourceName"     : "MyDocument",
      "section"        : "Introduction",
      "content"        : "Azure AI Search is a fully managed cloud search service...",
      "embedding"      : [ 0.0023064255, -0.009327292, ... ]  // 1536 floats
    },
    // more chunks...
  ]
}
08

Error Responses

200
OK — Full Azure AI Search payload returned.
401
Unauthorized — api-key header missing or wrong.
400
Bad Request — Required header missing, tenant_id/document_id non-numeric, no file, or not a PDF. Also returned when performEmbedding: true but embedding headers are missing.
422
Unprocessable — No text extracted (scanned/image-only PDF).
502
Bad Gateway — Azure OpenAI embedding call failed. Includes the chunk index and upstream error.
405
Method Not Allowed — Request was not POST.
09

Next Step in RAG Pipeline

Step 2 — Index the response in Azure AI Search Take the full JSON response body from this API and POST it as-is to the Index Document API (Step 2):
POST https://<service>.search.windows.net/indexes/<index-name>/docs/index?api-version=2025-05-01-preview
with header api-key: <your-search-api-key> and Content-Type: application/json.