Elastic Inference Service
Hosted inference endpoint that runs Elastic-managed LLMs, the ELSER sparse-embedding model and third-party embedding models for ingest, search and chat without provisioning ML nodes in a customer's own Elasticsearch deployment.
Find alternatives to Elastic Inference Service
Billed per million tokens for every model on the service, on top of the Elastic Cloud subscription it is reached from. No marketing page exists — the docs name it and state the billing unit, admitted under the docs-plus-price ruling of 2026-09-04.
What it does
- Model inference
- Model hosting
- Vector search
How it compares
Sources
- Pricing
-
Official documentation · 4 Sep 2026
"All models on EIS incur a charge per million tokens … EIS is billed per million tokens used". A rate card per model is not printed on this page.
- Description
-
Official documentation · 4 Sep 2026
Docs: EIS lets you "use machine learning models for ingest, search, and chat independently of your Elasticsearch infrastructure"; lists Elastic Managed LLMs, ELSER on EIS and jina-embeddings-v3.
- Sold within
-
Official documentation · 29 Sep 2026
The docs call it "a global service on Elastic Cloud", metered separately from ML-node VCUs: "All models on EIS incur a charge per million tokens."
Also from Elastic
-
Elastic AI Assistant Platform
A conversational assistant embedded in Kibana that answers natural-language questions against a customer's own indexed data across Elastic's Observability, Security and Search solutions.
-
Elastic Agent Builder Agent platform
A builder for custom AI agents that answer questions and take actions over data indexed in Elasticsearch, using configurable tools, skills and prompts.
-
Elastic AI SOC Engine Application
An AI security-operations layer that correlates alerts from a customer's existing security tools, prioritizes threats and guides response workflows without replacing their SIEM.
-
Elastic Workflows Agent platform
An automation engine that runs both scripted steps and AI agents which reason through investigations and execute response actions against data in Elasticsearch.
-
Elastic AIOps Application
GenAI- and ML-driven capability inside Elastic Observability that automatically detects, diagnoses and helps resolve operational issues, providing recommended actions for SREs.
-
Elastic LLM Observability Application
Monitoring capability inside Elastic Observability for generative-AI and agentic applications: performance, cost control, guardrail tracking and reliability for GenAI workloads.
-
Elastic Security Application
Agentic security-operations platform unifying SIEM, XDR and native automation, with autonomous agents handling detection-to-response workflows and purpose-built AI skills for threat hunting, alert analysis and detection engineering; supports multiple LLMs including on-premises models.
Something wrong here? Send a correction — quote this product id: elastic-inference-service.