Transforming a CLI into an AI-Native CLI

AI / CLI / ESKit

August 2026 · ESKit experiment

Building an AI CLI Helper from the CLI's Own Argument Definitions

An experiment with argparse, Claude, and ESKit

I've been playing around with AI agents lately, and one thing I've been thinking about is how to add AI capabilities to ESKit, a small Elasticsearch CLI toolkit I built.

My first experiment was closer to a traditional AI agent: Claude gets a set of tools, the tools call ESKit functions, and the agent can eventually perform operations. I built a simulator so I could experiment with the architecture without immediately modifying the real ESKit. ESKit-AI

That worked surprisingly well. But after playing with it for a while, I started wondering about something simpler.

What if I just wanted to ask my CLI a question?

For example:

How do I create an index named test?

And have it answer:

eskit index create test

What does the AI actually need to know?

ESKit already has a pretty complete definition of its CLI. It's built with Python's argparse, so the parser knows things like:

  • commands and subcommands
  • arguments
  • flags
  • defaults
  • required arguments
  • choices
  • argument types
  • descriptions
  • help text

I initially considered creating another schema specifically for the AI. Then I thought:

What if I make argparse the source of truth?

I really don't like maintaining multiple definitions that describe essentially the same interface.

So I built a small function that walks the argparse parser and turns it into JSON.


{
    "program": "eskit index create",
    "description": "Create an Elasticsearch index.",
    "arguments": [
        {
            "flags": [],
            "name": "name",
            "type": "str",
            "required": true,
            "description": "Name of the index."
        },
        {
            "flags": ["--dry-run"],
            "name": "dry_run",
            "type": "boolean",
            "required": false,
            "default": false,
            "description": "Preview the operation without executing it."
        }
    ]
}
                

The JSON is basically a normalized description of the existing CLI.

And then I gave it to Claude

The first experiment was surprisingly effective.

eskit ai "How can I create an index named a?"

Claude could respond with:

eskit index create a

It could also understand optional arguments. For example, if I asked about creating a repository with a filesystem location, it could figure out:

eskit repo create backup --location /data/es

without me writing a separate AI-specific description of that command.

Then I started asking more interesting questions.

How can I rename an Elasticsearch index using ESKit?

There isn't a dedicated rename command. Claude figured out that ESKit can accomplish this by using reindex to copy the data to a new index and then deleting the original index.

That's where this started feeling less like "AI-powered help text" and more like an AI actually understanding the CLI.

Metadata Makes a Surprisingly Big Difference

I started wondering how much the additional metadata actually mattered, so I ran a simple A/B experiment. I gave Claude the same command description, once without ESKit-specific metadata and once with it.

Without metadata

For the question:

What operations can modify data?

Claude could still figure out quite a lot. It identified commands such as:


eskit repo create
eskit repo delete
eskit snap create
eskit snap delete
eskit snap restore
eskit index create
eskit index delete
eskit reindex
eskit archive pull
eskit archive sync
eskit archive push

But it didn't have a consistent classification. It grouped everything under "Destructive Operations", including commands such as snap create and index create.

With metadata

The same question produced a much more structured answer:

Write Operations (risk: write)
eskit pull
eskit repo create
eskit snap create
eskit snap restore
eskit index create
eskit reindex
eskit archive pull
eskit archive push

Destructive Operations (risk: destructive)
eskit repo delete
eskit snap delete
eskit index delete
eskit archive sync

That's already a much better representation of the model's understanding of the CLI.

But the more interesting test was:

Can I delete a snapshot without user confirmation?

Without metadata, Claude answered:

Yes, you can delete a snapshot without an interactive confirmation prompt...

It then explained that --force could override safety checks.

That answer isn't completely unreasonable if you only look at the CLI arguments. The argparse definition tells the model that snap delete exists and that it has --force, --push, and --dry-run.

But it doesn't explicitly tell the model that user confirmation is a required safety concept.

So I added:

{
    "metadata": {
        "risk": "destructive",
        "requires_confirmation": true,
        "description": "The user must explicitly confirm the action before execution."
    }
}

I asked exactly the same question again.

This time Claude answered:

No, you cannot delete a snapshot without user confirmation.

It also correctly distinguished --force from confirmation:

--force is an administrative override for safety checks and is not a substitute for user confirmation.

That was probably the most interesting result of the experiment.

Metadata isn't just documentation

This made me rethink what I mean by "command description."

The argparse information tells Claude what the CLI looks like.

The metadata tells Claude what the commands mean operationally.

argparse
snap delete
├── name
├── --dry-run
├── --push
└── --force
ESKit metadata
snap delete
├── risk: destructive
└── requires_confirmation: true

The second one isn't really needed to execute the CLI. It's there because I'm asking an AI to reason about the CLI.

And that seems to make a surprisingly large difference.

I'm sure there are more sophisticated ways to represent this information, but I like that the first version is so small. A few semantic fields produced noticeably different behavior from the model.

That's a pretty good return for a few extra fields.

A More Complex Question

One of the more interesting tests was asking Claude to go beyond individual commands and compose a workflow:

I have Elasticsearch indices named index-<date>. Can you give me a script to create a daily snapshot and transfer the data to my local host using ESKit?

Claude produced a surprisingly complete shell script. It inferred the date-based index name and composed several ESKit commands.

#!/bin/bash

REPO="my-repo"
ARCHIVE="my-archive"

TODAY=$(date +%Y-%m-%d)
INDEX="index-${TODAY}"
SNAPSHOT="snapshot-${TODAY}"
SNAPSHOT_NAME="${REPO}/${SNAPSHOT}"

eskit pull es

eskit snap create "${SNAPSHOT_NAME}" \
--index "${INDEX}" \
--wait

eskit pull archive

eskit archive push "${ARCHIVE}" 
--dst "localhost:/path/to/local/destination"

That's much more than a simple command lookup.

But there was an important mistake: it used archive push for the final transfer.

In ESKit, archive pull transfers archive data from the remote source to the local archive:

flowchart TD A["Remote source"] -->|"archive pull"| B["Local archive"]

archive push goes in the opposite direction:

flowchart TD A["Local archive"] -->|"archive push"| B["Remote destination"]

So the generated workflow was conceptually close, but the transfer direction was wrong.

I actually found this mistake useful. It showed me that the model can reason across multiple CLI commands and compose a workflow, but it also exposed a limitation of relying primarily on CLI argument definitions.

argparse tells Claude what commands and arguments exist. It doesn't necessarily tell it all of the semantic relationships between those commands.

That gives me some obvious areas to explore later: better descriptions, richer metadata, examples, and perhaps explicit relationships between commands.

There Is Still a Lot to Improve

The current implementation is deliberately simple. The complete command description is currently around 55 KB.

That's not enormous, but I don't want to blindly send the entire CLI description to the model for every question.

flowchart TD A["Argparse JSON"] --> B["Normalized/refined JSON"] B --> C["Projection"] C --> D["Context-specific command JSON"] D --> E["Claude"]

For example, a read-only helper might only receive commands marked risk: read. An operator might receive additional write operations. An agent could eventually have access to a much larger capability set.

There is also the obvious token-cost question. Sending the whole command tree on every request won't scale well as the CLI grows.

I haven't optimized that yet.

For now, I'm more interested in seeing how far the simple version can go.

Helper First, Agent Later

This experiment also changed how I think about the AI architecture for ESKit.

An AI helper doesn't necessarily need to execute anything. It can simply answer:

  • "How do I do X?"
  • "What operations are available?"
  • "Is this operation destructive?"

That's useful by itself. The user still has ultimate control over whether to execute the command.

But that doesn't mean AI shouldn't execute anything. It's simply another layer of AI integration that a CLI tool can have with relatively little effort.

It's a CLI-first approach.

The next experiment is to use the same command description to turn CLI commands into LLM tool definitions, then translate the resulting tool call back into an argparse command.

flowchart TD A["Argparse"] --> B["Command JSON"] B --> C["Claude Tool Definition"] C --> D["Claude"] D --> E["index_create(name=test)"] E --> F["CLI adapter"] F --> G["args[index, create, test]"] G --> H["Argparse"] H --> I["Existing ESKit command"]

If that works, the AI could potentially use the existing CLI as its execution interface, rather than requiring a completely separate AI-specific execution layer.

That's a different architecture from a standalone AI agent, and I'm curious to see where the boundary is.

Why I'm Playing With This

It started as a way to learn more about LLMs and AI agents.

ESKit was a good candidate because I wanted to see what it would be like to interact with a real CLI using natural language.

Then it turned into a series of engineering questions.

I already had a CLI.
I already had argparse.
I already had descriptions and argument definitions.

So instead of building another abstraction immediately, I wanted to see what would happen if I simply exposed what I already had to an LLM.

So far, the CLI helper works pretty well. And it's surprisingly pleasant to interact with.

That's what makes this experiment fun.

I'm going to keep playing with it and see where the simple approach starts to break down.

The next experiment is probably going to be the interesting one: Can I turn this CLI helper into an actual CLI-centric AI agent?

If you're interested in the experiment, the ESKit branch is here:

ESKit AI implementation →
AI evaluation results →

The current implementation is intentionally small: a helper to dump argparse into JSON, some additional command metadata, and a simple AI helper.

This is part of an ongoing series of experiments exploring AI-assisted CLI interfaces and agent architectures in ESKit.

Next: Turning Tool Calls Back into Commands →