Chunk and Perform Embedding
Accepts a PDF binary, splits the extracted text into structured chunks (800–1200 chars each), optionally calls the Azure OpenAI Embeddings endpoint per chunk, and returns a payload you can POST directly to Azure AI Search — no modification needed.
| Property | Value |
|---|---|
| Method | POST |
| Content-Type | multipart/form-data |
| All parameters | Passed as request headers (not query params or form fields) |
| File | Binary PDF in the document form field (multipart body) |
| Response | Exact Azure AI Search Documents – Index API payload ({ "value": [...] }) |
Lab Credentials — For Training Use Only
| Header | Type | Required | Description |
|---|---|---|---|
| api-key | String | Required | Fixed lab API key. Returns 401 if missing or wrong. |
| Header | Type | Required | Description |
|---|---|---|---|
| sourceName | String | Required | Logical name of the source document (e.g. AzureGuide). Written to every chunk's sourceName field. |
| tenant-id | Int64 | Required | Tenant identifier. Must be numeric. Written to every chunk's tenant_id field. |
| document-id | Int64 | Required | Document identifier. Must be numeric. Written to every chunk's document_id field. |
| Header | Type | Required | Description |
|---|---|---|---|
| performEmbedding | Boolean | Optional | Pass true to call Azure OpenAI and populate embedding on every chunk. Defaults to false — embedding is [] when omitted. |
| embeddingEndpoint | String (URL) | Conditional | Your Azure OpenAI base URL, e.g. https://my-resource.openai.azure.com. Required when performEmbedding: true. |
| embeddingKey | String | Conditional | Azure OpenAI API key for your resource. Required when performEmbedding: true. |
| embeddingDeploymentName | String | Conditional | Name of the embedding deployment (e.g. text-embedding-3-small). Required when performEmbedding: true. |
{embeddingEndpoint}/openai/deployments/{embeddingDeploymentName}/embeddings?api-version=2024-02-01content is sent as the input field. The returned 1536-dimension float array is placed directly into the chunk's embedding field.
If any chunk fails, the entire request returns a 502 error with the chunk index and reason.
| Field | Type | Required | Description |
|---|---|---|---|
| document | Binary (PDF) | Required | The PDF file. Must contain selectable text (not a scanned/image-only PDF). Send as multipart/form-data. |
cURL — without embedding (embedding will be [])
curl -X POST \ -H "api-key: ankitg.in" \ -H "sourceName: AzureGuide" \ -H "tenant-id: 1" \ -H "document-id: 42" \ -H "performEmbedding: false" \ -F "document=@/path/to/document.pdf" \ https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php
cURL — with embedding (embedding array populated with 1536 floats)
curl -X POST \ -H "api-key: ankitg.in" \ -H "sourceName: AzureGuide" \ -H "tenant-id: 1" \ -H "document-id: 42" \ -H "performEmbedding: true" \ -H "embeddingEndpoint: https://my-resource.openai.azure.com" \ -H "embeddingKey: <your-azure-openai-key>" \ -H "embeddingDeploymentName: text-embedding-3-small" \ -F "document=@/path/to/document.pdf" \ https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php
JavaScript (fetch) — with embedding
const form = new FormData(); form.append('document', pdfFile); // File from <input type="file"> const res = await fetch('https://labs.lowcademy.com/apis/chunk-and-perform-embedding.php', { method : 'POST', headers: { 'api-key' : 'ankitg.in', 'sourceName' : 'AzureGuide', 'tenant-id' : '1', 'document-id' : '42', 'performEmbedding' : 'true', 'embeddingEndpoint' : 'https://my-resource.openai.azure.com', 'embeddingKey' : '<your-azure-openai-key>', 'embeddingDeploymentName' : 'text-embedding-3-small' }, body: form }); const payload = await res.json(); // payload.value is ready to POST to Azure AI Search
OutSystems — Consume REST API (OnBeforeRequest)
// Set all parameters as request headers in the OnBeforeRequest handler: Request.Headers.Add("api-key", "ankitg.in") Request.Headers.Add("sourceName", TextVar_SourceName) Request.Headers.Add("tenant-id", IntToText(TenantId)) Request.Headers.Add("document-id", IntToText(DocumentId)) Request.Headers.Add("performEmbedding", "true") Request.Headers.Add("embeddingEndpoint", TextVar_EmbeddingEndpoint) Request.Headers.Add("embeddingKey", TextVar_EmbeddingKey) Request.Headers.Add("embeddingDeploymentName", TextVar_DeploymentName) // Send the PDF binary as multipart/form-data body field: document
POST https://<service>.search.windows.net/indexes/<index>/docs/index?api-version=2024-07-01value array with @search.action: "upload" is the complete payload.
| Field | Type | Description |
|---|---|---|
| value | Array | Array of document objects, ready for Azure AI Search indexing. |
| value[]@search.action | String | Always "upload". |
| value[].id | String | Unique key: {tenant_id}-{document_id}-{zero-padded index}, e.g. 1-42-0001. |
| value[].tenant_id | Int64 | Echoed from the tenant_id header. |
| value[].document_id | Int64 | Echoed from the document_id header. |
| value[].sourceName | String | Echoed from the sourceName header. |
| value[].section | String | Detected heading for this chunk (max 90 chars). |
| value[].content | String | Chunk text, 800–1200 characters. |
| value[].embedding | Array<float> | 1536-dimension float array when performEmbedding: true. Empty array [] otherwise. |
Sample Response — performEmbedding: false
{
"value": [
{
"@search.action" : "upload",
"id" : "1-42-0001",
"tenant_id" : 1,
"document_id" : 42,
"sourceName" : "AzureGuide",
"section" : "Introduction",
"content" : "Azure AI Search is a fully managed cloud search service...",
"embedding" : []
}
]
}Sample Response — performEmbedding: true
{
"value": [
{
"@search.action" : "upload",
"id" : "1-42-0001",
"tenant_id" : 1,
"document_id" : 42,
"sourceName" : "AzureGuide",
"section" : "Introduction",
"content" : "Azure AI Search is a fully managed cloud search service...",
"embedding" : [ 0.0023064255, -0.009327292, 0.015797119, /* ... 1536 values total */ ]
}
]
}api-key header missing or wrong.{ "success": false, "error": "Unauthorized. Invalid or missing api-key header." }tenant_id / document_id is non-numeric, file field missing, or file is not a PDF.performEmbedding: true but any embedding header is absent.{ "success": false, "error": "Header embeddingEndpoint is required when performEmbedding is true." }{ "success": false, "error": "No text could be extracted from the PDF." }{ "success": false, "error": "Embedding failed on chunk 3: Access denied. Check your embeddingKey." }| Behaviour | Detail |
|---|---|
| Chunk size | 800–1200 characters. Breaks at paragraph or sentence boundaries — never mid-word. |
| Section detection | Numbered headings (e.g. 1.2 Overview), ALL-CAPS short lines, or standard keywords (Chapter, Section, Introduction…). Max 90 chars. |
| TOC filtering | Table-of-contents pages are automatically excluded. |
| Chunk ID | {tenant_id}-{document_id}-{zero-padded-index}, e.g. 1-42-0003. |
| Embedding model | Calls whichever deployment you specify. Use text-embedding-3-small (1536 dims) to match the index schema. |
| Embedding calls | Sequential — one cURL call per chunk, 30 s timeout each. For large PDFs (many chunks) this adds latency proportional to chunk count. |
| Embedding failure | Fails fast on the first error and returns a 502 with the chunk index and upstream message. |
| Response metadata | Chunk count available in the X-Total-Chunks response header. Embedding status in X-Embedding-Performed. |
| Text extraction | Uses pdftotext (Poppler) when available, PHP fallback for uncompressed streams. |
| CORS | Access-Control-Allow-Origin: * |
POST https://<service>.search.windows.net/indexes/<index-name>/docs/index?api-version=2024-07-01api-key: <your-search-api-key> and Content-Type: application/json.