← Back to Sample APIs

Chunk and Perform Embedding

Accepts a PDF binary, splits the extracted text into structured chunks (800–1200 chars each), optionally calls the Azure OpenAI Embeddings endpoint per chunk, and returns a payload you can POST directly to Azure AI Search — no modification needed.

Concept: File Upload · All-Header Auth Azure OpenAI Embeddings · Azure AI Search
01

Endpoint

POST https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php
PropertyValue
MethodPOST
Content-Typemultipart/form-data
All parametersPassed as request headers (not query params or form fields)
FileBinary PDF in the document form field (multipart body)
ResponseExact Azure AI Search Documents – Index API payload ({ "value": [...] })
02

Authentication Header

Lab Credentials — For Training Use Only

api-keyankitg.in
HeaderTypeRequiredDescription
api-key String Required Fixed lab API key. Returns 401 if missing or wrong.
03

Document Headers (always required)

HeaderTypeRequiredDescription
sourceName String Required Logical name of the source document (e.g. AzureGuide). Written to every chunk's sourceName field.
tenant-id Int64 Required Tenant identifier. Must be numeric. Written to every chunk's tenant_id field.
document-id Int64 Required Document identifier. Must be numeric. Written to every chunk's document_id field.
04

Embedding Headers (conditional)

HeaderTypeRequiredDescription
performEmbedding Boolean Optional Pass true to call Azure OpenAI and populate embedding on every chunk. Defaults to false — embedding is [] when omitted.
embeddingEndpoint String (URL) Conditional Your Azure OpenAI base URL, e.g. https://my-resource.openai.azure.com. Required when performEmbedding: true.
embeddingKey String Conditional Azure OpenAI API key for your resource. Required when performEmbedding: true.
embeddingDeploymentName String Conditional Name of the embedding deployment (e.g. text-embedding-3-small). Required when performEmbedding: true.
How embedding is called The API constructs the Azure OpenAI URL as:
{embeddingEndpoint}/openai/deployments/{embeddingDeploymentName}/embeddings?api-version=2024-02-01

Each chunk's content is sent as the input field. The returned 1536-dimension float array is placed directly into the chunk's embedding field. If any chunk fails, the entire request returns a 502 error with the chunk index and reason.
05

Request Body — File Upload

FieldTypeRequiredDescription
document Binary (PDF) Required The PDF file. Must contain selectable text (not a scanned/image-only PDF). Send as multipart/form-data.
06

Sample Requests

cURL — without embedding (embedding will be [])

curl -X POST \
  -H "api-key: ankitg.in" \
  -H "sourceName: AzureGuide" \
  -H "tenant-id: 1" \
  -H "document-id: 42" \
  -H "performEmbedding: false" \
  -F "document=@/path/to/document.pdf" \
  https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php

cURL — with embedding (embedding array populated with 1536 floats)

curl -X POST \
  -H "api-key: ankitg.in" \
  -H "sourceName: AzureGuide" \
  -H "tenant-id: 1" \
  -H "document-id: 42" \
  -H "performEmbedding: true" \
  -H "embeddingEndpoint: https://my-resource.openai.azure.com" \
  -H "embeddingKey: <your-azure-openai-key>" \
  -H "embeddingDeploymentName: text-embedding-3-small" \
  -F "document=@/path/to/document.pdf" \
  https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php

JavaScript (fetch) — with embedding

const form = new FormData();
form.append('document', pdfFile); // File from <input type="file">

const res = await fetch('https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php', {
  method : 'POST',
  headers: {
    'api-key'                  : 'ankitg.in',
    'sourceName'               : 'AzureGuide',
    'tenant-id'                : '1',
    'document-id'              : '42',
    'performEmbedding'         : 'true',
    'embeddingEndpoint'        : 'https://my-resource.openai.azure.com',
    'embeddingKey'             : '<your-azure-openai-key>',
    'embeddingDeploymentName'  : 'text-embedding-3-small'
  },
  body: form
});
const payload = await res.json();
// payload.value is ready to POST to Azure AI Search

OutSystems — Consume REST API (OnBeforeRequest)

// Set all parameters as request headers in the OnBeforeRequest handler:
Request.Headers.Add("api-key",                 "ankitg.in")
Request.Headers.Add("sourceName",               TextVar_SourceName)
Request.Headers.Add("tenant-id",                IntToText(TenantId))
Request.Headers.Add("document-id",              IntToText(DocumentId))
Request.Headers.Add("performEmbedding",         "true")
Request.Headers.Add("embeddingEndpoint",        TextVar_EmbeddingEndpoint)
Request.Headers.Add("embeddingKey",             TextVar_EmbeddingKey)
Request.Headers.Add("embeddingDeploymentName",  TextVar_DeploymentName)

// Send the PDF binary as multipart/form-data body field: document
07

Response — 200 OK (Azure AI Search Payload)

Direct copy-paste to Azure AI Search The response body is exactly the JSON you POST to:
POST https://<service>.search.windows.net/indexes/<index>/docs/index?api-version=2024-07-01

No wrapping, no modification. The value array with @search.action: "upload" is the complete payload.
FieldTypeDescription
valueArrayArray of document objects, ready for Azure AI Search indexing.
value[]@search.actionStringAlways "upload".
value[].idStringUnique key: {tenant_id}-{document_id}-{zero-padded index}, e.g. 1-42-0001.
value[].tenant_idInt64Echoed from the tenant_id header.
value[].document_idInt64Echoed from the document_id header.
value[].sourceNameStringEchoed from the sourceName header.
value[].sectionStringDetected heading for this chunk (max 90 chars).
value[].contentStringChunk text, 800–1200 characters.
value[].embeddingArray<float>1536-dimension float array when performEmbedding: true. Empty array [] otherwise.

Sample Response — performEmbedding: false

{
  "value": [
    {
      "@search.action"  : "upload",
      "id"              : "1-42-0001",
      "tenant_id"       : 1,
      "document_id"     : 42,
      "sourceName"      : "AzureGuide",
      "section"         : "Introduction",
      "content"         : "Azure AI Search is a fully managed cloud search service...",
      "embedding"       : []
    }
  ]
}

Sample Response — performEmbedding: true

{
  "value": [
    {
      "@search.action"  : "upload",
      "id"              : "1-42-0001",
      "tenant_id"       : 1,
      "document_id"     : 42,
      "sourceName"      : "AzureGuide",
      "section"         : "Introduction",
      "content"         : "Azure AI Search is a fully managed cloud search service...",
      "embedding"       : [ 0.0023064255, -0.009327292, 0.015797119, /* ... 1536 values total */ ]
    }
  ]
}
08

Error Responses

200
OK — Full Azure AI Search payload returned.
401
Unauthorized — api-key header missing or wrong.
{ "success": false, "error": "Unauthorized. Invalid or missing api-key header." }
400
Bad Request — A required header is missing, tenant_id / document_id is non-numeric, file field missing, or file is not a PDF.
Also returned when performEmbedding: true but any embedding header is absent.
{ "success": false, "error": "Header embeddingEndpoint is required when performEmbedding is true." }
422
Unprocessable Entity — No text could be extracted (scanned / image-only PDF).
{ "success": false, "error": "No text could be extracted from the PDF." }
502
Bad Gateway — Azure OpenAI embedding call failed. Includes which chunk failed and the upstream error message.
{ "success": false, "error": "Embedding failed on chunk 3: Access denied. Check your embeddingKey." }
405
Method Not Allowed — Request method was not POST.
09

Behaviour Notes

BehaviourDetail
Chunk size800–1200 characters. Breaks at paragraph or sentence boundaries — never mid-word.
Section detectionNumbered headings (e.g. 1.2 Overview), ALL-CAPS short lines, or standard keywords (Chapter, Section, Introduction…). Max 90 chars.
TOC filteringTable-of-contents pages are automatically excluded.
Chunk ID{tenant_id}-{document_id}-{zero-padded-index}, e.g. 1-42-0003.
Embedding modelCalls whichever deployment you specify. Use text-embedding-3-small (1536 dims) to match the index schema.
Embedding callsSequential — one cURL call per chunk, 30 s timeout each. For large PDFs (many chunks) this adds latency proportional to chunk count.
Embedding failureFails fast on the first error and returns a 502 with the chunk index and upstream message.
Response metadataChunk count available in the X-Total-Chunks response header. Embedding status in X-Embedding-Performed.
Text extractionUses pdftotext (Poppler) when available, PHP fallback for uncompressed streams.
CORSAccess-Control-Allow-Origin: *
Tip — Sending to Azure AI Search Copy the entire response body and POST it as-is to:
POST https://<service>.search.windows.net/indexes/<index-name>/docs/index?api-version=2024-07-01
with header api-key: <your-search-api-key> and Content-Type: application/json.