Previous: ← From 25K to 7K Tokens: Simplifying the CLI Definition Without Losing Context

Transforming a CLI into an AI-Native CLI

AI / CLI / ESKit

August 2026 · ESKit experiment

Agent Loop: Letting the LLM Decide When to Act, Ask, and Stop

Moving from one-shot tool calls to a multi-step agent

From One Tool Call to an Agent

In the previous parts of this series, I built the pieces needed to let an LLM interact with ESKit. The CLI's own argparse definitions became the source for LLM tool definitions, tool calls could be translated back into ESKit commands, execution and presentation were separated, and the CLI definition was simplified from roughly 25K tokens to about 7K tokens.

At that point, the LLM could call an ESKit operation. But it wasn't really an agent yet. The missing piece was the agent loop.

From One Tool Call to Multiple Steps

The first version of the AI integration was essentially a one-shot interaction: the user made a request, the LLM selected a tool, ESKit executed it, and the result was returned to the LLM.

flowchart TD A["User"] --> B["LLM"] B -->|"Tool Call"| C["ESKit"] C -->|"Tool Result"| B B -->|"Final Response"| A

That works well when the request maps cleanly to a single operation. But many useful requests don't. For example, a user might ask: "Can you get the latest cache and show me available indices sorted by document count?"

There isn't necessarily one ESKit operation that directly answers that question. The agent needs to get the latest cache, inspect the result, retrieve the relevant index information, sort it, and then present the answer.

flowchart TD A["User Request"] --> B["LLM"] B -->|"pull cache"| C["ESKit"] C -->|"Cache Result"| B B -->|"get index information"| C C -->|"Index Result"| B B --> E["Analyze + Sort"] E --> F["Final Response"]

This is where the agent loop becomes useful. After receiving a tool result, the LLM can decide whether it needs another tool call or whether it has enough information to answer the user.

The ask_user Tool

Once the agent can perform multiple operations, another problem appears. An LLM can ask the user questions, which means the application needs a way to distinguish a request for input from a final response. I didn't want to interpret the generated text to figure out whether the LLM was asking a question or simply making a statement.

That's why I tried to add an ask_user tool for this purpose. The tool gives the LLM a way to request input from the user before continuing, and I can decide that it's the end of the interaction when there are no more tool calls.

For example, if I first create an index and then ask the agent to delete it, the interaction can look like this:

$ eskit ai "Can you delete test-index-99?"

LLM → ask_user
      "This will delete index test-index-99. Continue?"

User → yes

LLM → index_delete
      {"index": "test-index-99", "force": true}

ESKit → SUCCESS

LLM → "Index deleted successfully."

The important part is that the application doesn't automatically enter an interactive confirmation mode every time a tool is called. Instead, the LLM decides whether it needs to ask the user something.

Here's another example of how this can play out. I asked Claude to create an index, but initially didn't provide a name. The agent asked for the missing information and then asked whether a mapping should be applied. After I answered both questions, Claude produced the corresponding ESKit command but did not actually execute it. The important part is that I did not make the CLI to become interactive unconditionally when using the AI-mode.

$ eskit ai "Can you create an index?"

LLM → ask_user
      "What is the name of the index you want to create?"

User → test-index99

LLM → ask_user
      "Would you like to apply a specific mapping?"

User → no

LLM → Final Response
      "eskit index create test-index99"

Please note that Claude decided that presenting the command was an appropriate final response for this particular interaction. The same agent is also capable of taking the next step and calling a tool to actually create the index when it determines that execution is appropriate. This is one of the things I like about this agent loop architecture: I don't have to force every interaction into an execution path or give it a system prompt to instruct it to use tools all the time.

The LLM can ask for information, call a tool, or simply return a response. When it does choose to act, the request still goes through ESKit's normal validation, permissions, and safety mechanisms. And if I want to, I can also ask the agent to create the index for me explicitly, and it will do so.

Don't Force Every Interaction to Be Interactive

One approach would have been to make every agent request interactive, requiring the user to approve each step. I didn't want to do that. A question such as "Why are some indices yellow?" doesn't require user input. The agent can inspect the available information and explain the result on its own.

Likewise, a request such as:

eskit ai "Take a look at Host X and give me a quick health report. Tell me anything that looks worth investigating."

is much more useful if the agent can investigate on its own. It can decide which tools to use, inspect their results, and continue until it has enough information to provide a useful report.

This means the agent has three possible directions after each step:

flowchart TD A["LLM"] -->|"Tool Call"| B["ESKit"] A -->|"ask_user"| C["User Input"] A -->|"Final Response"| D["User"]

The LLM decides whether it needs another tool, needs the user's input, or has enough information to finish.

A Health Report Example

Now, taking look at the an example of question above to let agent to investigate the health of my HP-DO environment:

eskit ai "Take a look at HP-DO and give me a quick health report. Tell me anything that looks worth investigating."

This request is intentionally vague. There isn't a predefined checklist in the application telling the agent exactly which commands it must run. Instead, the agent can inspect the information available through ESKit and decide what is relevant.

In one run, the agent noticed that a number of Metricbeat indices were yellow and explained that this can happen when primary shards are allocated but replica shards cannot be allocated. On a single-node Elasticsearch deployment, that can be expected if replicas are configured but there isn't another node available to hold them.

It also identified other things worth investigating, including snapshot currency and the state of individual indices.

This is one of the parts of the experiment I find particularly interesting. The value isn't necessarily that the LLM knows something that ESKit doesn't. The value is that it can combine several pieces of information and decide what deserves attention.

Output from Claude:

$ eskit ai "Take a look at HP-DO and give me a quick health report. Tell me anything that looks worth investigating."

## HP-DO Health Report

### Overall Status: ⚠️ YELLOW — Attention Needed

**Cluster:** tpotcluster (Elasticsearch 9.3.5)
**Last Updated:** 2026-08-31 03:20 UTC

### Critical Issues

**Yellow Index Health**
- 31 indices are reporting yellow health.
- All 31 Metricbeat indices are affected.
- The likely cause is replica shards that cannot be allocated on the current cluster.

### Data Overview

- **67** active indices
- **36** green
- **31** yellow
- Largest index: `logstash-2026.08.30` — 1 GB / 2.7M documents
- Recent data is flowing through September 1.
- All indices are managed by the `tpot` ILM policy.

### Backup Status

- Repository: `daily_repo`
- Latest snapshot found: **2026-08-23**
- Recent snapshots report `SUCCESS` with no failures.
- Data is current through **2026-09-01**, so the snapshot date is worth investigating.

### Worth Investigating

1. **Yellow replicas** — Check node count and shard allocation.
2. **Snapshot age** — Verify that daily snapshots are still being created.
3. **Test index** — `test_logstash-v2-test` appears to be a leftover test index.
4. **Large index** — `logstash-2026.08.30` is significantly larger than surrounding days.

Would you like me to investigate any of these issues in more detail?

                

Coming back to the ask_user tool, the LLM asking a question in its final response, and there is an important distinction between requesting input to continue and offering to do more work.

In this example, the agent finishes its health report with:

Would you like me to investigate any of these issues in more detail?

This does not trigger the ask_user tool. The agent has already completed the task, so there is nothing it needs from the user to finish. It is simply offering a possible next step. The ask_user tool is different. It is used when the agent needs user input before it can continue—for example, when it wants to perform a potentially destructive operation and needs confirmation.

I like this distinction because it means the application doesn't have to decide in advance whether an interaction should be conversational or interactive. The LLM can complete a task on its own, ask the user for input when it is actually necessary, or finish with an invitation to continue.

The agent loop does not mean that every request must eventually result in a tool call. The LLM decides whether it has enough information to respond, ask the user for additional information, or take an action through ESKit. In some cases, as shown below, the agent may gather the information needed to operate but choose to present the corresponding ESKit command instead of executing it. I don't consider this a problem—the goal is not to force the LLM to execute every possible action, but to give it the ability to act when appropriate while keeping execution and safety within ESKit.

flowchart TD A["User"] -->|Request/Answer|B["LLM"] B -->|"Need more information"| A B -->|"Tool Call"| D["ESKit"] D -->|"Tool Result"| B B -->|"Final Response
e.g. show command"| A

Other notable instructions I gave the LLM include:

  • Please get the latest cache, and show me available indices sorted by document counts.
  • Which indices are expiring within 7 days?
  • Are there any unhealthy indices?

The interesting part is that there is no specific functionality in ESKit to sort or filter indices by expiration date or status. Claude was able to use information returned from multiple ESKit commands or filter the results based on the question. In addition, it used the output in human-friendly formats such as tables and sectional summaries instead of structured data like JSON.

The Agent Can Also Be Wrong

Of course, this is still an LLM. During testing, I asked about the yellow indices and what could be done to fix them. The model correctly identified the general cause: replica shards could not be allocated. However, when it moved from diagnosis to remediation, it suggested an ESKit command that didn't actually exist for changing the replica count. That was a useful failure.

The important thing is that neither ESKit nor the LLM could execute the invalid command. ESKit does not know such a command, and the LLM does not have the authority to execute arbitrary commands. The correct Elasticsearch operation would involve changing the index settings, such as setting the replica count to zero on a single-node deployment. But ESKit did not expose that operation as an AI tool, so the agent could not perform it through ESKit which is expected.

In addition, the agent doesn't have any idea about how to contact the host or Elasticsearch cluster. It doesn't have access to any credentials to either SSH or Elasticsearch. It's all handled by ESKit.

This reinforced one of the architectural principles from the earlier parts of the series:

The LLM should not be the authority for execution.

It can reason, investigate, and make recommendations, but ESKit remains responsible for validating and executing operations.

The Boundary Still Matters

With the agent loop in place, the overall architecture now looks like this:

flowchart TD A["User"] -->|Request/Answer| B["LLM"] B -->|"Tool Call"| C["ESKit"] C -->|"Tool Result"| B B -->|"ask_user"| A C --> E["Validation"] E --> F["Permissions / Safety"] F --> G["Execution"] G --> C B -->|"Final Response"| A

The LLM controls the reasoning loop. ESKit controls what can actually happen. That separation becomes more important as the agent becomes more capable. From a security perspective, I think it's important to keep a clear boundary between User, LLM, and ESKit.

Also, during the development of the agent loop, I didn't worry about the LLM making a mistake and executing something destructive because the LLM doesn't have the authority to do it. Even if it uses a tool call, ESKit still asks the user for confirmation. There is an option to force the execution for administrative uses, but if I want to, I can hide it from the LLM tools. So if someone says to skip confirmation and execute the tool, it's okay for the LLM to issue the tool call because ESKit will enforce the confirmation.

On top of that, ESKit has a concept of a "protected" host, where configuration must be set to execute any modifying operations, including deletions. This configuration can be managed per ESKit user.

What Changed?

The implementation itself wasn't much code. The conceptual change was more important. Before the agent loop, the application essentially said: "Here is a tool. Call it."

Now it says: "Here are the tools. Decide what you need to do, use them as necessary, ask the user when you need input, and stop when you have completed the task." That small change turns a tool-calling interface into something much closer to an agent. And importantly, I didn't have to build a complicated planner. The LLM became the planner for the reasoning loop, while ESKit remained the authority for execution.

Learnings So Far

A few things stood out from these experiments.

Multi-step tasks are where the agent becomes useful

Single tool calls are interesting, but the ability to inspect a result and decide what to do next is where the real value starts to appear.

User interaction doesn't have to be hard-coded

Rather than forcing every operation into an interactive confirmation flow, ask_user gives the LLM a mechanism for requesting input only when it actually needs it.

The LLM should not be trusted as the execution authority

It can reason, investigate, and make recommendations, but ESKit remains responsible for validating and executing operations.

Hallucinations are still useful to observe

The incorrect recommendation for changing shard settings was a good reminder that a convincing explanation isn't necessarily an executable one. The architecture needs to assume that the model will occasionally be wrong.

The CLI is becoming an agent interface without replacing the CLI

The original ESKit commands are still there. The AI layer provides another way to interact with them—one that can reason across multiple commands and present the result more naturally.

From CLI Helper to AI-Native CLI

Looking back at the progression so far:

flowchart TD A["Part 1 - Argparse → AI CLI Helper"] A --> C["Part 2 - Tool calls → ESKit commands"] C --> D["Part 3 - Separate Execution and Presentation"] D --> E["Part 4 - 25K tokens → 7K-token"] E --> F["Part 5 - One tool call → Agent loop"]

The original goal was fairly simple: make ESKit easier to use through natural language. The interesting part has been discovering that the existing CLI already contains much of the structure needed for an AI interface.

The AI layer doesn't need to replace the CLI. It can sit on top of it.

And now that the agent loop is working, the next question becomes less about whether the LLM can operate ESKit, and more about how well we can observe, control, and improve that operation.

ESKit AI implementation →

This is part of an ongoing series of experiments exploring AI-assisted CLI interfaces and agent architectures in ESKit.

Next: Observing the Agent: Adding Events and Tracing to ESKit AI →