Transforming a CLI into an AI-Native CLI
AI / CLI / ESKit
August 2026 · ESKit experiment
Turning Tool Calls Back into Commands
In the previous entries, I experimented with giving Claude the
command definitions from an existing CLI and using them as the
context for an AI helper.
The idea was deliberately simple: ESKit already had
argparse, so instead of creating another AI-specific
description of the CLI, I could turn the existing definitions
into JSON and give that to the model.
That worked well enough that I started wondering about the next step.
What if Claude could actually execute the commands?
This time I wanted to keep the same basic philosophy. I didn't want to build a completely separate execution layer just for the AI. The goal was to let the LLM produce a structured tool call, translate that tool call back into normal CLI arguments, and then let the existing ESKit command execute it.
Turning CLI Commands into Tools
The first part was relatively straightforward. The command definitions I already had could be transformed into LLM tool definitions.
For example, the existing index create command
could become a tool with an argument such as
index.
I could then ask:
eskit ai "Can you create an index called test-index?"
Instead of Claude simply returning text describing the command, it produced a structured tool call:
ToolUseBlock(
name='index_create',
input={
'index': 'test-index'
})
That's an important change. The model wasn't just generating a command string. It was producing structured data that my program could inspect.
Tool Calls Can Become CLI Arguments
The next question was how to get from that structured tool call back to the existing CLI.
I introduced a small ToolCall representation:
@dataclass
class ToolCall:
id: str
name: str
arguments: dict
The adapter takes the tool name and arguments and finds the corresponding command definition.
For example:
ToolCall(
name="repo_show",
arguments={
"name": "daily_repo"
}
)
becomes:
["repo", "show", "daily_repo"]
And:
ToolCall(
name="cat",
arguments={
"kind": "index",
"json": True
}
)
becomes:
["cat", "index", "--json"]
Finally, those arguments can be passed to the same
argparse parser that a normal user would use.
This was the part I was particularly interested in. The AI doesn't need a second implementation of the command. It can use the CLI itself as the execution interface.
A Small Detail: Positional Arguments
While implementing the adapter, I ran into a small but important detail. Not every argument has flags. A positional argument might be represented in the command definition like this:
{
"flags": [],
"name": "name",
"type": "str",
"required": true
}
While an optional argument might look like:
{
"flags": ["--json"],
"name": "json",
"type": "boolean",
"required": false,
"default": false
}
So the adapter needs to know whether an argument is positional or optional.
I initially considered preserving all of the argparse flags in the AI-facing definition, but that creates another interesting optimization opportunity: the model doesn't necessarily need every representation detail that the CLI needs.
The AI mostly needs to understand the argument's name, type, meaning, choices, and whether it is required. The CLI adapter can retain the knowledge of how that argument maps to actual flags.
The First Real Tool Calls
Once the conversion was working, I started trying ordinary ESKit operations through the AI interface.
A simple read operation worked:
eskit ai "Can you show me a detail of repository called daily_repo?"
Claude produced:
ToolUseBlock(
name='repo_show',
input={
'name': 'daily_repo'
})
args: ['repo', 'show', 'daily_repo']
Another example was asking for cached index information:
eskit ai "Can you show me the index from cache?"
Claude selected the cat tool and supplied the
appropriate positional value:
input={'kind': 'index'}
args: ['cat', 'index']
It also understood optional output formatting:
eskit ai "Can you show list of index from cache in json format?"
input={'kind': 'index', 'json': True}
args: ['cat', 'index', '--json']
At this point the proof of concept felt surprisingly solid.
It Can Execute Real Commands
The next test was an operation that actually modifies Elasticsearch.
eskit ai "Can you create ai index called test-index2?"
Claude produced a tool call equivalent to:
{'index': 'test-index2'}
which became:
['index', 'create', 'test-index2']
And ESKit executed it:
Index created successfully.
I also tested a snapshot operation:
eskit ai "Can you create a snapshot named 2026.08.22 in daily_repo repository?"
Claude represented the repository and snapshot name using the argument expected by the existing command:
{'name': 'daily_repo/2026.08.22'}
which the adapter converted into:
['snap', 'create', 'daily_repo/2026.08.22']
This was a nice milestone for the experiment. The AI wasn't just explaining ESKit anymore. It could use the existing CLI to perform an operation.
Then I Tried Multiple Operations
This is where the difference between a tool call and an actual agent started becoming obvious.
I asked Claude to perform several operations:
eskit ai "Can you set a host to HP-Local, pull latest cache, show me index from the cache?"
Claude actually returned multiple tool calls:
ToolUseBlock(
name='host_set',
input={
'host': 'HP-Local'
}
)
ToolUseBlock(
name='pull',
input={
'host': 'HP-Local'
}
)
That was interesting because the model had clearly understood that there were multiple operations involved. But my implementation only executed the first tool call. That exposed the next architectural step.
A tool call isn't the same thing as an agent loop.
The model can request a tool, but if the application doesn't execute that tool, return the result to the model, and allow the model to continue, the interaction stops there. My current implementation only goes through the first part of that loop. In other words, I had built enough infrastructure for an agent, but I hadn't actually built the agent loop yet.
The Difference Between a Helper and an Agent
This distinction made the architecture much clearer to me.
A simple helper can do:
Or, with one executable operation:
An agent needs the loop:
That sounds obvious when written down, but seeing the behavior in a real CLI made the distinction much more concrete.
And Now Token Optimization Matters
The other thing that became much more obvious once tool execution started working was the size of the context. My command description was around 55 KB. The model was receiving essentially the entire CLI description for every request.
The actual tool calls were tiny. For example:
{"kind": "index"}
But the model had to reason about thousands of lines of
command definitions to produce that tiny result.
I started looking at the command JSON more critically.
There was a lot of information that was useful to
argparse but didn't necessarily need to be sent
to the LLM.
Optimization #1: Remove Redundant Information
The first optimization is intentionally boring: don't send information the model doesn't need.
For example, the command definition contains a lot of fields like:
{
"flags": [
"-v",
"--verbose"
],
"name": "verbose",
"type": "boolean",
"required": false,
"default": false,
"choices": null,
"nargs": 0,
"description": "Enable verbose logging"
}
Some of these fields are useful to the parser, but not all of them are useful to Claude.
In particular, values such as null are often
just noise.
If there are many arguments, repeatedly sending
"choices": null doesn't add much information.
Similarly, if an argument has the normal default for its
type, there may be no reason to explicitly send that default.
This is a simple optimization, but it is also the easiest one to apply because it doesn't change the semantics of anything.
It's basically:
Remove what doesn't carry useful information.
Optimization #1.5: Define Common Arguments Once
The next optimization is slightly more interesting.
ESKit has arguments that appear across many commands:
--verbose--debug--json--config--host--view
If the same definition appears in dozens of commands, sending the complete definition every time is redundant.
Instead, I can define the argument once:
"common_arguments": {
"verbose": {
"type": "boolean",
"description": "Enable verbose logging"
},
"debug": {
"type": "boolean",
"description": "Enable debug logging"
},
"json": {
"type": "boolean",
"description": "Output in JSON format"
}
}
And then reference it from commands:
"commands": {
"init": {
"arguments": [
"verbose",
"debug",
"json",
"demo"
]
},
"host show": {
"arguments": [
"config",
"host",
"verbose",
"debug",
"json"
]
}}
This is still a small change, but now the command
representation is starting to look less like a raw dump of
argparse and more like an intermediate
representation designed for another stage.
The Compiler Analogy
This is where I started thinking about the architecture differently.
The JSON originally started as a dump of the argparse structure. But it is increasingly becoming something closer to an intermediate representation.
There are effectively two compilation directions here.
The first is:
CLI → AI
The existing CLI definition is transformed into a representation the LLM can understand.
The second is:
AI → CLI
A structured tool call is transformed back into the argument list expected by the existing CLI.
I wasn't really thinking about this as a compiler when I started the project. It emerged naturally from trying to avoid duplicating the CLI implementation. And I like the analogy because it also gives the optimization work a useful name.
Removing redundant fields, deduplicating common arguments, projecting only relevant commands, and eventually generating role-specific tool definitions all start to look like intermediate representation optimization.
Safety and Permissions Start to Matter More
Another thing that became more important once the AI could
actually execute commands was the metadata I experimented
with in the previous post.
For a helper that only answers questions, metadata such as
risk and
requires_confirmation improves the model's
reasoning.
Once the model can execute commands, that metadata becomes
much more interesting.
The system can potentially use it to decide which tools an agent is allowed to see and which operations require confirmation.
Tool definition
│
├── risk: read
├── risk: write
└── risk: destructive
│
▼
Agent profile
│
▼
Available tools
For example, a read-only agent shouldn't necessarily receive destructive tools at all. That is different from telling the model: "please don't use this tool." The tool can simply not exist from that agent's perspective. That's a much more interesting direction for ESKit because it connects the AI experiment with a security concept: capability-based access to operations.
What I Have Now
At this point, the proof of concept can do more than I originally expected.
It can take an existing ESKit command definition, expose commands to Claude as tools, receive structured tool calls, translate those calls back into normal CLI arguments, and execute the existing ESKit command.
Some examples are now as simple as:
eskit ai "Can you pull cache?"
eskit ai "Can you set a host to HP-Local?"
eskit ai "Can you show me a detail of repository called daily_repo?"
eskit ai "Can you create ai index called test-index2?"
The AI can also combine operations conceptually and return multiple tool calls. My current implementation just doesn't complete the full loop yet.
That makes the next step pretty clear.
What's Next?
The next milestone is to complete the agent loop. But before doing that, I want to spend some time reducing the amount of information being sent to the model. The command definition started at roughly 55 KB. There is no reason to assume that the LLM needs all of that information for every request.
I want to explore a few simple optimizations first:
- remove redundant or empty properties
- omit information that can be inferred
- define common arguments once and reference them
- project only the relevant part of the command tree
- send only the tools needed for the current context
After that, the full agent loop should be a much more interesting experiment.
From a Simple Helper to an Agent
The interesting thing about this experiment is that I haven't really added much AI-specific functionality to ESKit. Most of the work has been adapting information that already existed.
The AI is effectively sitting between two representations of the same interface. That's what makes the approach appealing to me. I don't need to create an entirely new AI command system. I can keep improving the translation layer around the existing CLI.
And now the project has a fairly natural progression: first an AI helper, then executable tool calls, then a complete agent loop, and eventually permission-aware agents. None of those steps requires ESKit to stop being a normal CLI.
The next experiment: complete the tool-call loop and see what happens when Claude can execute a workflow instead of just one command.
If you're interested in the experiment, the ESKit branch is here:
This is part of an ongoing series of experiments exploring AI-assisted CLI interfaces and agent architectures in ESKit.