Transforming a CLI into an AI-Native CLI
AI / CLI / ESKit / Observability
September 2026 · ESKit experiment
Observing the Agent: Adding Events and Tracing to ESKit AI
The Agent Was Working — But What Was It Doing?
In the previous part of this series, I moved ESKit's AI interface from a one-shot tool call to a multi-step agent loop.
That changed the behavior of the application quite a bit. The LLM could now inspect information, call a tool, receive the result, decide what to do next, wait, ask the user for information, call another tool, and eventually decide that the task was complete.
This worked, but it introduced a new problem: the more autonomous the agent became, the harder it was to understand exactly what it had done.
A simple terminal output such as:
→ Using snap_create
→ Using wait_tool
→ Using snap_create
tells me what tools were called, but not the complete execution flow. I wanted to know which model response caused each tool call, what arguments were supplied, what results came back, and how the interaction progressed.
I also wanted to be able to inspect the execution after it happened.
Why Logging Wasn't Quite Enough
My first instinct was to add more logging to the agent loop. That would have been easy enough.
But the agent was already starting to have several different kinds of activity:
- LLM prompts
- LLM responses
- Tool calls
- Tool results
- User input
- Final responses
- Errors
I didn't want the agent implementation to become filled with
print() statements and logging code for every new event.
More importantly, I wanted the same execution information to be usable
by more than one consumer.
For example, the terminal might want a short progress message while a
trace file might want the complete structured event.
That led me to an event-based design.
Introducing an Event Bus
Instead of having the agent directly print information, the agent emits events.
The agent doesn't need to know what happens to an event after it is emitted. It simply reports what happened. The Event Bus then sends the event to any registered listeners.
This creates a useful separation: console output and persistent tracing become two different consumers of the same execution information.
Event Structure
I wanted the events to be more than arbitrary log messages. Each event has a type, timestamp, flow ID, sequence ID, and associated data.
{
"type_id": 1004,
"type": "ai.tool_call",
"timestamp": "2026-09-06T04:12:31.123456+00:00",
"flow_id": "675b3272446d42c18bea0f8edd72de6e",
"sequence_id": 7,
"data": {
"tool": "snap_create",
"arguments": {
"snapshot": "daily_repo/2026.09.01"
}
}
}
- The
type_idprovides a stable machine-readable identifier. - The
typemakes the event easier to inspect as a human. - The
flow_ididentifies the overall AI interaction. - The
sequence_idgives events their order within that flow.
I deliberately kept the flow information as metadata on the event rather than putting it into every event's data payload.
Defining the Events
The initial event set is intentionally small.
ai.run_started
ai.llm_prompt
ai.llm_response
ai.tool_call
ai.tool_result
ai.user_input_requested
ai.user_input_received
ai.function_call_started
ai.function_call_completed
ai.final_response
ai.error
Event Flow:
These events describe the major transitions in an AI interaction without trying to model every internal detail of the application. I also created an event registry and generated the Python event definitions from it. This gives each event a stable numeric identifier while keeping the human-readable event name in one place.
The generated emitter gives the agent code simple methods such as:
emitter.tool_call(
tool="snap_create",
arguments={
"snapshot": "daily_repo/2026.09.01"
}
)
emitter.tool_result(
tool="snap_create",
result=result
)
The agent doesn't need to construct the event object or manage sequence numbers itself. The Event Bus handles that part.
Why a Sequence ID?
The sequence ID is a small detail, but it became useful almost immediately.
An agent interaction can produce many events:
1 ai.run_started
2 ai.llm_prompt
3 ai.llm_response
4 ai.tool_call
5 ai.function_call_started
6 ai.function_call_completed
7 ai.tool_result
8 ai.llm_prompt
9 ai.llm_response
10 ai.tool_call
11 ai.user_input_requested
12 ai.user_input_received
13 ai.tool_result
...
Timestamps are useful, but a sequence number gives me an explicit ordering within the flow. It also means different event types don't need to maintain their own counters. The Event Bus owns the ordering for the entire interaction. That makes it easier to reconstruct what happened later.
Separating Console Output from Execution
Once events existed, I moved the AI-specific terminal output into a dedicated listener.
For example, a normal run might display:
AI flow started (claude-haiku-4-5-20251001)
→ Using ask_user
← Received response from: ask_user
→ Using index_create
← Received response from: index_create
→ Using wait_tool
← Received response from: wait_tool
AI final response received
With verbose output enabled, the listener can show additional information such as tool arguments, flow IDs, and execution timing. The important part is that none of this presentation logic needs to be embedded in the agent loop itself.
The agent reports events. The listener decides what is appropriate to display. This also means I can change the terminal presentation without changing the agent's behavior.
Adding Persistent Tracing
The next step was to make the same events persistent. I added an AI tracer that writes each event as a JSON object on its own line.
The result is a JSONL trace:
{"type_id":1001,"type":"ai.run_started",...}
{"type_id":1002,"type":"ai.llm_prompt",...}
{"type_id":1003,"type":"ai.llm_response",...}
{"type_id":1004,"type":"ai.tool_call",...}
{"type_id":1005,"type":"ai.tool_result",...}
{"type_id":1002,"type":"ai.llm_prompt",...}
...
Tracing is optional and can be enabled with:
eskit ai "show me the latest snapshots" --trace
Trace files are automatically created under:
.eskit/traces/
The file name includes the timestamp and part of the flow ID, which makes individual executions easy to identify.
Why JSONL?
I considered the trace simply as a JSON document containing an array of events, but JSONL was a better fit for how the tracer works. Events can be written as they happen rather than waiting until the entire interaction has completed. That has a useful property for an agent: if the process fails halfway through, the trace still contains the events that were successfully written before the failure. JSONL also makes the trace easy to process later. Individual events can be read line by line without loading the entire execution into memory.
It also leaves the door open for future analysis, replay, or importing execution events into another system.
The Trace Changed How I Debug the Agent
When I repeated one of the snapshot prompts after adding the tracer, I noticed something interesting. I ran essentially the same request to create snapshots and wait about 30 seconds between operations.
eskit ai "Can you create missing snapshots ... Please wait for each snapshot to be created, and before creating the next snapshot, please wait about 30 seconds."
In one run, the agent used ask_user as a way to pause before
continuing.
→ Using ask_user
Waiting 30 seconds before creating snapshot ...
> ok
← User Input Received
The tracing and terminal output made that behavior visible, but it also made me think about a broader question: what happens when I want the agent to perform an action while still retaining deterministic control over part of that behavior?
This led me to introduce core tools for operations where I wanted
the LLM to request the action, but I wanted to control how the action
actually happened.
For this experiment, I introduced wait_tool.
The tool accepts the previous tool and next tool as part of its arguments,
along with the requested duration.
After adding wait_tool, the agent used it in the next run with
the same prompt.
→ Using wait_tool {'previous_tool': 'index_create', 'next_tool': 'index_delete', 'duration': 5}
That was useful by itself, but the trace made the behavior much easier to understand.
Instead of only seeing:
→ Using snap_create
→ Using wait_tool
→ Using snap_create
I could inspect the sequence of model responses and tool calls that led to those actions. This matters because LLM behavior is not deterministic in the same way as a traditional function call.
If an agent behaves differently between two runs, I want an artifact that lets me compare the executions rather than relying on terminal output or memory. The trace effectively becomes an execution history for the agent.
Observability Becomes More Important as the Agent Grows
This made me think about the difference between a normal CLI operation and an agent operation.
A traditional command might have a relatively simple execution path:
The agent introduces another layer:
The execution path is now dynamic. The LLM can decide whether to call another tool, ask the user something, or stop. That makes observability less like traditional logging and more like recording an execution flow.
The Event System Is Not Just for AI
Although I built the first events for the AI interface, I intentionally kept the Event Bus generic. The event itself doesn't know that it is being used for AI. It contains a type, timestamp, metadata, and event data.
That means the same infrastructure could eventually be used for other ESKit operations if there is a useful reason to observe them. For example, a future listener could send selected operational events somewhere else without changing the code that performs the operation. I don't have a specific need for that yet, so I'm deliberately leaving the system small.
For now, the useful abstraction is simply:
The code reports what happened; listeners decide what to do with it.
Tracing Also Introduces a New Responsibility
Recording everything creates another problem: the trace can contain a lot
of information.
An ai.llm_prompt event can contain the system prompt,
conversation messages, and tool definitions.
A tool result can also contain information that may not be appropriate
to persist without filtering.
During development I also encountered an important reminder that debugging output can contain sensitive configuration information. That means the tracer cannot simply be treated as a harmless dump of everything the application knows. For the initial implementation, I wanted to establish the event and trace pipeline first. The next step is to introduce a deliberate serialization and sanitization boundary so sensitive values can be filtered before they are written to persistent traces. I also don't want the trace to become a recording of hidden model reasoning. The goal is to capture the observable execution: prompts, responses, tool calls, results, user interaction, and metadata useful for debugging and analysis.
Learnings
-
Humans need an execution history too
Once the LLM can take multiple actions, a final response doesn't tell me enough about what happened along the way. The sequence of events is part of the useful information.
-
Events are cleaner than scattered logging
The Event Bus gives the agent one simple responsibility: report events. Console output, tracing, and future consumers can be implemented independently.
-
JSONL is a useful format for execution traces
Events can be persisted incrementally, inspected easily, and processed later without requiring a large in-memory execution record.
-
Non-deterministic behavior needs observability
When the same request can result in different tool choices or execution paths, having a persistent trace makes those differences observable.
-
If a deterministic behavior is required, make a tool
When I want the agent to perform an action but need deterministic control over how that action happens, I can provide a tool that represents the action. The LLM still decides when to request the tool, while ESKit controls the actual behavior.
-
Tracing needs a security boundary
An execution trace can contain much more information than a normal log. That makes sanitization and careful handling of persistent traces important as the system evolves.
From Agent to Observable Agent
The progression of the series now looks like this:
At this point, ESKit has moved quite a bit beyond the original experiment of simply asking an LLM to generate CLI commands. The AI layer can now reason across multiple operations, interact with the user, execute tools through ESKit, and produce a persistent record of what happened. That last part is becoming increasingly important. As the agent becomes more capable, I don't just want it to be able to act. I want to understand how it acted.
The next question is therefore not simply how to give the agent more capabilities, but how to give it the right context at the right time without sending everything every time. And, as the system grows, I want to continue refining the boundary between the LLM and ESKit.
This is part of an ongoing series of experiments exploring AI-assisted CLI interfaces and agent architectures in ESKit.