Onboard search data
Use Elasticsearch APIs to onboard application search data.
Typical workloads include:
- application records;
- product or content catalogs;
- documents;
- knowledge-base content;
- full-text search;
- semantic search;
- vector search;
- retrieval-augmented generation data.
Unlike observability telemetry, search documents do not need to follow the OpenTelemetry data model.
Recommended architecture
Application / Data Pipeline
↓
Elasticsearch API
↓
Index or data stream
↓
Elasticsearch
↓
Search / Kibana / AI
Use the Elasticsearch endpoint provided for your deployment.
Do not construct endpoints from Azure resource names.
Before you begin
Make sure:
- network access to Elasticsearch is working;
- DNS and TLS validation succeed;
- the application has a scoped credential;
- the target index or data-stream design has been defined;
- mappings have been reviewed for the intended search workload.
Do not use the administrator account for normal application ingestion.
Design the data model
Before loading production data, decide how the application data should be represented.
Consider:
- document identity;
- searchable text fields;
- exact-match keyword fields;
- dates and numeric fields;
- nested or object fields;
- language requirements;
- vector fields when semantic search is required;
- retention and deletion requirements.
Avoid relying on unrestricted dynamic mapping for important production schemas.
Index naming
Use names that clearly identify the application and purpose.
Examples:
products
knowledge-base
customer-documents
support-content
For time-oriented application data, a data stream may be more appropriate.
For persistent application search content, a regular Elasticsearch index is often simpler.
Create mappings
Define mappings before bulk-loading important production data.
Example:
{
"mappings": {
"properties": {
"title": {
"type": "text"
},
"body": {
"type": "text"
},
"source_id": {
"type": "keyword"
},
"updated_at": {
"type": "date"
}
}
}
}
The appropriate mapping depends on the application's query patterns.
Application credentials
Create a dedicated API key or application identity with only the privileges required for the target indices.
A typical ingestion application needs permission to write to its own indices but should not have unrestricted access to the cluster or unrelated data.
Store the credential in the application's approved secrets-management system.
Do not embed administrator credentials in:
- source code;
- container images;
- configuration repositories;
- documentation.
Index documents
Applications can send documents directly through the Elasticsearch API.
Example:
curl --fail-with-body \
--request POST \
"https://<elasticsearch-hostname>/knowledge-base/_doc" \
--header "Authorization: ApiKey <application-api-key>" \
--header "Content-Type: application/json" \
--data '{
"title": "Example document",
"body": "Content that will be indexed for search.",
"source_id": "doc-001",
"updated_at": "2026-08-18T00:00:00Z"
}'
For production workloads, applications should use the appropriate Elasticsearch client library rather than relying on individual curl requests.
Bulk ingestion
For large initial data loads or high-volume indexing, use the Elasticsearch Bulk API or an appropriate client-library bulk helper.
Batching reduces request overhead and usually provides better throughput than sending one document per request.
Monitor:
- indexing errors;
- rejected requests;
- response latency;
- document counts;
- storage growth.
Do not increase batch sizes without monitoring the impact on the platform.
Vector and semantic search
For semantic or AI-enabled search, documents can include vector representations in addition to normal searchable fields.
A typical document can contain:
Document
├─ Metadata
├─ Searchable text
└─ Vector representation
The exact vector mapping and embedding model depend on the application's search design.
Keep the original document identity and metadata so search results can be traced back to their source.
Validate the first data
Before loading the complete production dataset:
- index a representative set of documents;
- confirm document counts;
- run representative search queries;
- verify text analysis and mappings;
- confirm filtering and sorting work as expected;
- validate relevance;
- validate vector or semantic search when used;
- check indexing performance and storage growth.
Fix mapping or analysis problems before loading a large dataset.
Handle updates and deletes
Applications should have a clear strategy for keeping Elasticsearch synchronized with the source system.
Common patterns include:
- update documents using a stable document ID;
- reindex documents when source content changes;
- delete documents when the source record is removed;
- rebuild an index when major schema changes are required.
Avoid creating duplicates when a source document is updated.
Production considerations
Before scaling the workload, review:
- expected document count;
- indexing rate;
- average document size;
- query rate;
- retention;
- replica requirements;
- mapping size;
- vector dimensions where used.
Monitor actual usage after onboarding and adjust platform capacity when required.
Need help
For assistance with onboarding or platform configuration:
- AI chat:
https://copilot.opsflw.io - Support portal:
https://support.ivedha.com/
Provide the deployment reference and a description of the workload. Do not include API keys, passwords, or sensitive production documents.
Next step
After data is indexed and validated, continue with search and visualize data.