Previous: ← Building an AI CLI Helper from the CLI's Own Argument Definitions

Transforming a CLI into an AI-Native CLI

AI / CLI / ESKit

August 2026 · ESKit experiment

Turning Tool Calls Back into Commands

Tool calls, argparse, agent loops, and the beginning of token optimization

In the previous entries, I experimented with giving Claude the command definitions from an existing CLI and using them as the context for an AI helper. The idea was deliberately simple: ESKit already had argparse, so instead of creating another AI-specific description of the CLI, I could turn the existing definitions into JSON and give that to the model.

That worked well enough that I started wondering about the next step.

What if Claude could actually execute the commands?

This time I wanted to keep the same basic philosophy. I didn't want to build a completely separate execution layer just for the AI. The goal was to let the LLM produce a structured tool call, translate that tool call back into normal CLI arguments, and then let the existing ESKit command execute it.

Turning CLI Commands into Tools

The first part was relatively straightforward. The command definitions I already had could be transformed into LLM tool definitions.

flowchart TD A["Argparse"] --> B["Command JSON"] B --> C["LLM Tool Definition"] C --> D["Claude"]

For example, the existing index create command could become a tool with an argument such as index.

I could then ask:

eskit ai "Can you create an index called test-index?"

Instead of Claude simply returning text describing the command, it produced a structured tool call:

ToolUseBlock(
name='index_create',
input={
    'index': 'test-index'
})

That's an important change. The model wasn't just generating a command string. It was producing structured data that my program could inspect.

Tool Calls Can Become CLI Arguments

The next question was how to get from that structured tool call back to the existing CLI.

I introduced a small ToolCall representation:

@dataclass
class ToolCall:
    id: str
    name: str
    arguments: dict

The adapter takes the tool name and arguments and finds the corresponding command definition.

For example:

ToolCall(
    name="repo_show",
    arguments={
        "name": "daily_repo"
    }
)

becomes:

["repo", "show", "daily_repo"]

And:

ToolCall(
    name="cat",
    arguments={
        "kind": "index",
        "json": True
    }
)

becomes:

["cat", "index", "--json"]

Finally, those arguments can be passed to the same argparse parser that a normal user would use.

flowchart TD A["Claude"] --> B["ToolCall"] B --> C["CLI adapter"] C --> D["args:[cat, index, --json]"] D --> E["Argparse"] E --> F["Existing ESkit command"]

This was the part I was particularly interested in. The AI doesn't need a second implementation of the command. It can use the CLI itself as the execution interface.

A Small Detail: Positional Arguments

While implementing the adapter, I ran into a small but important detail. Not every argument has flags. A positional argument might be represented in the command definition like this:

{
    "flags": [],
    "name": "name",
    "type": "str",
    "required": true
}

While an optional argument might look like:

{
    "flags": ["--json"],
    "name": "json",
    "type": "boolean",
    "required": false,
    "default": false
}

So the adapter needs to know whether an argument is positional or optional.

I initially considered preserving all of the argparse flags in the AI-facing definition, but that creates another interesting optimization opportunity: the model doesn't necessarily need every representation detail that the CLI needs.

The AI mostly needs to understand the argument's name, type, meaning, choices, and whether it is required. The CLI adapter can retain the knowledge of how that argument maps to actual flags.

The First Real Tool Calls

Once the conversion was working, I started trying ordinary ESKit operations through the AI interface.

A simple read operation worked:

eskit ai "Can you show me a detail of repository called daily_repo?"

Claude produced:

ToolUseBlock(
name='repo_show',
input={
    'name': 'daily_repo'
})

args: ['repo', 'show', 'daily_repo']

Another example was asking for cached index information:

eskit ai "Can you show me the index from cache?"

Claude selected the cat tool and supplied the appropriate positional value:

input={'kind': 'index'}

args: ['cat', 'index']

It also understood optional output formatting:

eskit ai "Can you show list of index from cache in json format?"
input={'kind': 'index', 'json': True}

args: ['cat', 'index', '--json']

At this point the proof of concept felt surprisingly solid.

It Can Execute Real Commands

The next test was an operation that actually modifies Elasticsearch.

eskit ai "Can you create ai index called test-index2?"

Claude produced a tool call equivalent to:

{'index': 'test-index2'}

which became:

['index', 'create', 'test-index2']

And ESKit executed it:

Index created successfully.

I also tested a snapshot operation:

eskit ai "Can you create a snapshot named 2026.08.22 in daily_repo repository?"

Claude represented the repository and snapshot name using the argument expected by the existing command:

{'name': 'daily_repo/2026.08.22'}

which the adapter converted into:

['snap', 'create', 'daily_repo/2026.08.22']

This was a nice milestone for the experiment. The AI wasn't just explaining ESKit anymore. It could use the existing CLI to perform an operation.

Then I Tried Multiple Operations

This is where the difference between a tool call and an actual agent started becoming obvious.

I asked Claude to perform several operations:

eskit ai "Can you set a host to HP-Local, pull latest cache, show me index from the cache?"

Claude actually returned multiple tool calls:


ToolUseBlock(
    name='host_set',
    input={
        'host': 'HP-Local'
    }
)

ToolUseBlock(
    name='pull',
    input={
        'host': 'HP-Local'
    }
)

That was interesting because the model had clearly understood that there were multiple operations involved. But my implementation only executed the first tool call. That exposed the next architectural step.

A tool call isn't the same thing as an agent loop.

The model can request a tool, but if the application doesn't execute that tool, return the result to the model, and allow the model to continue, the interaction stops there. My current implementation only goes through the first part of that loop. In other words, I had built enough infrastructure for an agent, but I hadn't actually built the agent loop yet.

The Difference Between a Helper and an Agent

This distinction made the architecture much clearer to me.

A simple helper can do:

flowchart TD A["Question"] --> B["Claude"] B --> C["Answer"]

Or, with one executable operation:

flowchart TD A["Question"] --> B["Claude"] B --> C["ToolCall"] C --> D["Execute"] D --> E["Exit"]

An agent needs the loop:

flowchart LR A["Question"] --> B["Claude"] B --> C["Tool call"] C --> D["Execute"] D --> E["Tool result"] E --> F{"More Tool Calls?"} F -->|yes| B F -->|no| G["Final Answer"]

That sounds obvious when written down, but seeing the behavior in a real CLI made the distinction much more concrete.

And Now Token Optimization Matters

The other thing that became much more obvious once tool execution started working was the size of the context. My command description was around 55 KB. The model was receiving essentially the entire CLI description for every request.

The actual tool calls were tiny. For example:

{"kind": "index"}

But the model had to reason about thousands of lines of command definitions to produce that tiny result. I started looking at the command JSON more critically. There was a lot of information that was useful to argparse but didn't necessarily need to be sent to the LLM.

Optimization #1: Remove Redundant Information

The first optimization is intentionally boring: don't send information the model doesn't need.

For example, the command definition contains a lot of fields like:

{
    "flags": [
        "-v",
        "--verbose"
    ],
    "name": "verbose",
    "type": "boolean",
    "required": false,
    "default": false,
    "choices": null,
    "nargs": 0,
    "description": "Enable verbose logging"
}

Some of these fields are useful to the parser, but not all of them are useful to Claude.

In particular, values such as null are often just noise. If there are many arguments, repeatedly sending "choices": null doesn't add much information. Similarly, if an argument has the normal default for its type, there may be no reason to explicitly send that default.

This is a simple optimization, but it is also the easiest one to apply because it doesn't change the semantics of anything.

It's basically:

Remove what doesn't carry useful information.

Optimization #1.5: Define Common Arguments Once

The next optimization is slightly more interesting.

ESKit has arguments that appear across many commands:

  • --verbose
  • --debug
  • --json
  • --config
  • --host
  • --view

If the same definition appears in dozens of commands, sending the complete definition every time is redundant.

Instead, I can define the argument once:

"common_arguments": {
    "verbose": {
        "type": "boolean",
        "description": "Enable verbose logging"
    },
    "debug": {
        "type": "boolean",
        "description": "Enable debug logging"
    },
    "json": {
        "type": "boolean",
        "description": "Output in JSON format"
    }
}

And then reference it from commands:

"commands": {
"init": {
    "arguments": [
        "verbose",
        "debug",
        "json",
        "demo"
    ]
},
"host show": {
    "arguments": [
        "config",
        "host",
        "verbose",
        "debug",
        "json"
    ]
}}

This is still a small change, but now the command representation is starting to look less like a raw dump of argparse and more like an intermediate representation designed for another stage.

The Compiler Analogy

This is where I started thinking about the architecture differently.

The JSON originally started as a dump of the argparse structure. But it is increasingly becoming something closer to an intermediate representation.

flowchart TD A["argparse"] B["Normalized command representation"] C["Optimization / projection"] D["Context-specific representation"] E["AI Helper"] F["LLM Tool Definitions"] G["ToolCall"] H["CLI Adapter"] I["argparse"] J["Existing ESKit"] A --> B B --> C C --> D D --> E D --> F E --> G F --> G G --> H H --> I I --> J

There are effectively two compilation directions here.

The first is:

CLI → AI

The existing CLI definition is transformed into a representation the LLM can understand.

The second is:

AI → CLI

A structured tool call is transformed back into the argument list expected by the existing CLI.

I wasn't really thinking about this as a compiler when I started the project. It emerged naturally from trying to avoid duplicating the CLI implementation. And I like the analogy because it also gives the optimization work a useful name.

Removing redundant fields, deduplicating common arguments, projecting only relevant commands, and eventually generating role-specific tool definitions all start to look like intermediate representation optimization.

Safety and Permissions Start to Matter More

Another thing that became more important once the AI could actually execute commands was the metadata I experimented with in the previous post. For a helper that only answers questions, metadata such as risk and requires_confirmation improves the model's reasoning. Once the model can execute commands, that metadata becomes much more interesting.

The system can potentially use it to decide which tools an agent is allowed to see and which operations require confirmation.

Tool definition
│
├── risk: read
├── risk: write
└── risk: destructive
        │
        ▼
  Agent profile
        │
        ▼
  Available tools

For example, a read-only agent shouldn't necessarily receive destructive tools at all. That is different from telling the model: "please don't use this tool." The tool can simply not exist from that agent's perspective. That's a much more interesting direction for ESKit because it connects the AI experiment with a security concept: capability-based access to operations.

What I Have Now

At this point, the proof of concept can do more than I originally expected.

It can take an existing ESKit command definition, expose commands to Claude as tools, receive structured tool calls, translate those calls back into normal CLI arguments, and execute the existing ESKit command.

Some examples are now as simple as:

eskit ai "Can you pull cache?"
eskit ai "Can you set a host to HP-Local?"
eskit ai "Can you show me a detail of repository called daily_repo?"
eskit ai "Can you create ai index called test-index2?"

The AI can also combine operations conceptually and return multiple tool calls. My current implementation just doesn't complete the full loop yet.

That makes the next step pretty clear.

What's Next?

The next milestone is to complete the agent loop. But before doing that, I want to spend some time reducing the amount of information being sent to the model. The command definition started at roughly 55 KB. There is no reason to assume that the LLM needs all of that information for every request.

I want to explore a few simple optimizations first:

  • remove redundant or empty properties
  • omit information that can be inferred
  • define common arguments once and reference them
  • project only the relevant part of the command tree
  • send only the tools needed for the current context

After that, the full agent loop should be a much more interesting experiment.

From a Simple Helper to an Agent

The interesting thing about this experiment is that I haven't really added much AI-specific functionality to ESKit. Most of the work has been adapting information that already existed.

The AI is effectively sitting between two representations of the same interface. That's what makes the approach appealing to me. I don't need to create an entirely new AI command system. I can keep improving the translation layer around the existing CLI.

And now the project has a fairly natural progression: first an AI helper, then executable tool calls, then a complete agent loop, and eventually permission-aware agents. None of those steps requires ESKit to stop being a normal CLI.

The next experiment: complete the tool-call loop and see what happens when Claude can execute a workflow instead of just one command.

If you're interested in the experiment, the ESKit branch is here:

ESKit AI implementation →

This is part of an ongoing series of experiments exploring AI-assisted CLI interfaces and agent architectures in ESKit.

Next: Separating Execution from Presentation in ESKit →