I asked my own plugin for the most expensive piece of clothing in an inventory index, and it gave me a backpack.
The index had a field called dept with three values: FW, AP and EQ. To you and me that is footwear, apparel and equipment. To a language model reading a mapping, it is three pairs of letters. Nothing in the mapping says AP means apparel, so the model guessed, picked the wrong code, and wrote a confident sentence about the result. The worst part is that a wrong guess looks exactly like a right one from the outside. You get documents, you get an answer, and nothing tells you the question was misread.
Fixing that changed how nlsearch works more than anything else I have done to it. This post walks through what the plugin does, how a request moves through it, how it learns what an index means, and what is still rough.
What nlsearch is
nlsearch is an Elasticsearch plugin I build on weekends. It adds one endpoint, _nl. You send it a sentence, it asks a language model to turn that sentence into the matching Elasticsearch call, runs the call, and gives you back the result along with what it did.
POST /_nl
{"prompt": "red shoes under 50"}
{
"model": "ollama/qwen2.5-coder:7b",
"action": "search",
"index": "products",
"body": {"query": {"bool": {
"must": [{"match": {"name": {"query": "red shoes", "operator": "and", "fuzziness": "AUTO"}}}],
"filter": [{"term": {"category": "shoes"}}, {"range": {"price": {"lt": 50}}}]
}}},
"result": {"took": 3, "hits": {"total": {"value": 2, "relation": "eq"}, "hits": ["..."]}}
}
Depending on the response mode you ask for, you also get a short plain-English answer next to the raw result. It can search, add one document or many, update and delete by id or by query, create and delete indices, read and change mappings, and list what is there.
I built it with chat bots in mind. There is a single endpoint, no index name in the URL, and an optional session string so follow-ups like "now only the ones in stock" work. Any string will do as a session: a Slack chat id, a row id, a UUID in the browser.
Day to day I work on Elasticsearch and big-data search, so most of my effort went into the Elasticsearch side being strict, and letting the model be the loose part.
How a request flows
The interesting code is in NLRestHandler, Planner, Actions, Indices and Analysis. A request to _nl goes like this:
- Read the
prompt, plus optionalsession,dry_runandresponse. - If there is a session, load its history from a hidden system index,
.nlsearch-history. - Load the mappings of every index, then the "facts": for small keyword fields, the actual values (a terms aggregation, kept only if there are 40 or fewer), and min and max for numbers and dates.
- Load the briefing for each index. If one is missing or stale, study the index on the spot and store the result. More on briefings below.
- Add one sample document per index, but only for indices that still have no briefing.
- If all of that is over about 12,000 characters, drop the samples first, then the facts, then the field lists.
- Build the messages: the system prompt, the last 10 turns of the conversation with what happened after each one, then the context and the request itself, last.
- Ask the model for a plan, a single JSON object.
- Check the plan, build a real Elasticsearch request from it, and run it, unless this is a dry run.
- If Elasticsearch rejects it, give the model the reason and let it try once more.
- Turn the result into JSON, ask the model for a two or three sentence explanation if the mode wants one, save the turn to history, and respond.
The order of the final user message matters more than I expected. A comment in Planner.java records one failed experiment: when I put the previous outcome after the request, the model started acting on the outcome instead of the question. The request now always goes last.
Dates get help too. Small models are bad at calendar arithmetic, so the plugin works out date ranges itself and hands them over, instead of hoping the model knows when last month started.
The plan format
The model has to answer with one JSON object and nothing else. The fields are action (required), index, id, body, docs for bulk writes, and text for replies. The allowed actions are:
search, index, bulk, update, delete, update_by_query, delete_by_query, create_index, delete_index, get_mapping, put_mapping, list_indices and reply.
Models do not always follow instructions, so Plan.parse() is forgiving. It pulls the object out of code fences or surrounding chatter by taking everything from the first { to the last }. Actions also tolerates badly nested output for documents and mappings. If it still cannot make sense of a plan, that is a 502, blamed on the model, not a 500.
A few rules from the prompt that took me a while to get right:
- Text search always uses a
matchwith"operator": "and"and"fuzziness": "AUTO", so plurals and typos still land. "shoes" finds "Shoe". - Keyword values that appear in the facts become
termfilters, with the spelling copied from the list. - Anything in quotes, of any style, is taken literally.
- "Cut", "drop" or "forget" in the middle of a chat means narrow the results, not delete data.
replyis allowed in only a few cases. One of them is when a question has two equally good readings across two fields, because a query against the wrong field does not fail. It comes back looking like an answer.
How it learns what an index means
Index briefings
This is the backpack fix. Before answering anything about an index, nlsearch now reads the data. For every value of every small keyword field, it pulls a couple of documents that actually carry that value, using a terms aggregation with a top_hits sub-aggregation. One model call turns that into a short briefing in prose, and the result goes into a new system index, .nlsearch-analysis.
For the inventory index, the briefing ends up saying things like "FW = footwear (shoes, boots, sandals), AP = apparel (clothing, jackets, jeans), EQ = equipment (packs, poles, stoves)". From then on "how much equipment do we stock" becomes {"term": {"dept": "EQ"}} instead of a guess.
You do not have to call anything. The first question about an index triggers the analysis, so only that first question pays for it. The briefing is rebuilt when the mapping changes (the plugin keeps a CRC32 fingerprint of it) or when nlsearch.analysis_ttl runs out, 24 hours by default. There are endpoints if you want control:
| Request | What it does |
|---|---|
POST /_nl/analyze | Study everything, or one index, and store the result |
POST /_nl/analyze?force=true | Redo it even if the stored briefing is still fresh |
GET /_nl/analyze | Show what is stored, without looking again |
Pass a session to /_nl/analyze and the briefing is kept for that conversation alone. That is how you correct it for one chat without changing what everyone else sees.
A couple of small decisions I like here. An empty index is skipped, so it does not get stored as "nothing to know" and stay that way after data arrives. An index whose values are already plain English gets a one-line placeholder briefing, so it is not studied again on every request. And the analysis prompt tells the model that the documents beat the letters of a code, and to say "could not tell" rather than settle for a plausible word.
It also made requests cheaper, which I did not expect. A briefing says in words what sample documents could only hint at, so an index with a briefing stops sending samples. A normal question now carries less context than it did before briefings existed.
Every rule in the prompt now says how to find an answer, never what the answer is. None of them names a field from any real dataset, because a rule that does only helps the one dataset with that field. Early on, the prompt had partly been taught my demo data by hand, and it showed the moment I pointed it at something else.
Smaller fixes that mattered
- Follow-ups stay on the same index. The planner now says "this conversation is about the index X" right before the request, so a follow-up does not wander off to another index.
- Better retry messages. When Elasticsearch rejects a plan, the model now gets the root cause, not just the outer parse error, and it is placed right before the request.
- Outcome wording. At first I summarised hits as
[id 1] Red Running Shoe. The model read that and started writing{"term": {"id": 1}}in its next plan. Hits are now described as "Red Running Shoe (document 1)", and that problem went away. - Analysis timeouts. The analysis gets the configured timeout once per index, and a timeout there now comes back as a 502 with advice instead of a 500.
- Cleaner code. Every cluster lookup moved into
Indices.java, so the REST handler went back to being about HTTP and the shape of a conversation.
Guardrails
Letting a model write delete requests needs some care. These are enforced in code, not just asked for in the prompt:
- GET is read-only. Over GET only
search,get_mapping,list_indicesandreplyrun. Anything else is refused. - One index for destructive actions.
delete_by_query,update_by_query,delete_indexandput_mappingonly run against one named index.*, comma lists and_allare refused, whatever the model says. - Refused is never retried. If the plugin refuses a plan, there is no second attempt, so a refused "delete everything" cannot quietly turn into a narrower delete.
- Its own indices are off limits. Plans that touch
.nlsearch-historyor.nlsearch-analysisare refused. - Validation before execution. Search bodies are parsed with Elasticsearch's real
SearchSourceBuilderparser before anything runs. - Dry runs. With
dry_runthe request is built and validated, so a bad query still fails, but it is not executed.
Some rules are prompt-only, like quoted literals and fuzziness, and the plugin runs with the caller's privileges. So the usual advice applies: give it a user with the permissions you are comfortable with, and dry-run destructive things.
Choosing a model
The models come through LangChain4j. Four providers are supported: Ollama, OpenAI, Anthropic and Gemini. The OpenAI one takes a base URL, so anything with an OpenAI-compatible API works too. The default is a small local model, qwen2.5-coder:7b on Ollama, because I wanted it to work on a laptop with no API key.
| Setting | Default | Notes |
|---|---|---|
nlsearch.provider | ollama | ollama, openai, anthropic or gemini |
nlsearch.model | qwen2.5-coder:7b | the model name as the provider knows it |
nlsearch.url | provider default | any OpenAI-compatible server for openai |
nlsearch.api_key | empty | hidden from the settings APIs |
nlsearch.timeout | 60s | how long to wait for the model |
nlsearch.analysis_ttl | 24h | how long a briefing stays fresh |
All of them can go in elasticsearch.yml or be changed on a running cluster with PUT _cluster/settings. Models are rebuilt lazily, so a bad API key fails a request instead of a node start. Temperature is 0 everywhere, and the client-side retries are turned off, because retrying happens at the plan level where the model can see what went wrong.
Testing and releasing
There are over 80 JUnit tests, and none of them need a running cluster or a model. They cover plan parsing, turning plans into Elasticsearch requests, the guardrails, the planner's message building, the model settings and the history code. Being able to run them all in seconds is the reason I kept them as pure unit tests.
The build is Gradle. ./gradlew bundle produces the plugin zip. GitHub Actions builds and tests every push on every branch, and a version tag publishes a release with the zip. The plugin is tied to the exact Elasticsearch version, so each Elasticsearch release needs its own build.
For poking at it by hand, the repo has around 50 ready-made requests as both a Bruno and a Postman collection, including a group that walks through a whole conversation and the stored history.
What is still rough
- Model calls block a thread on the node's
genericpool for as long as the call takes. Tens of concurrent chats per node are fine. Thousands are not. - Turns of one session should be sent one after another. Two at the same time can lose a turn of history.
- History is kept forever unless you clean it up. Only the last 10 turns are replayed, but nothing expires on its own. The collections show how to do it with plain Elasticsearch calls.
- The model reads field names and sample values from your data. A document that contains instructions can try to steer it, the same as any other prompt input.
- Each request makes two model calls, one to plan and one to explain, unless you ask for the raw result only.
Try it
You need JDK 21, a model, and the Elasticsearch version a release was built for. The plugin is tied to an exact Elasticsearch version, so pick the zip that matches yours from the latest release. The quickest model is Ollama:
ollama pull qwen2.5-coder:7b
bin/elasticsearch-plugin install <url of the zip from the releases page>
Restart the node, then:
curl -XPOST localhost:9200/_nl -H 'Content-Type: application/json' \
-d '{"prompt": "how many products are there per category"}'
The code is on GitHub under Apache 2.0, and the docs are at sheikmohammedsha.github.io/nlsearch. If you point it at your own data and it gets something wrong, I would like to hear about it. Wrong answers are how most of these fixes happened.
Comments
Post a Comment