Documentation
Zeian docs.
Zeian is in alpha. The surface described here is implemented locally and evolving quickly — where a feature is still in progress, the page says so.
Quickstart
Install
$ npm install -g @zeian/cli
$ zeian --version
zeian 0.4.2 (alpha)Initialize a suite
Run zeian init inside the repository that holds your agent. It creates suites/ with a starter task and a zeian.config.ts pointing at your agent entrypoint.
$ cd my-agent && zeian init
created suites/smoke.yaml
created zeian.config.tsRun
$ zeian run suites/smoke.yaml
✓ greets-and-routes pass · 3 calls · $0.008
1/1 passA green run means every task reached its expected outcome and its trajectory stayed inside the declared budget.
Concepts
Suites and tasks
A suite is a YAML file describing a family of realistic tasks. Each task declares an input, the tools the agent may use, and what success means. Suites live in your repo — they are code, reviewed like code.
Trajectory
The ordered record of everything the agent did: tool calls with arguments, retries, model turns, cost, and latency. Zeian asserts on the trajectory, not just the final answer — an agent that calls a refund twice and guesses its way to the right output is a production incident, not a pass.
Verdict
Per-task result: pass, fail, or regression. A regression means the task passed but its behavior drifted from the stored baseline — more calls, a different tool order, higher cost.
Baseline
A committed golden trajectory per task. Baselines are what make diff review possible: every run is compared against them and the drift is shown in the report.
Suite format
An annotated suite:
suite: booking-flow # suite id
agent: ./agents/travel # agent entrypoint
tools: [flights, bookings, refunds] # MCP servers allowed
tasks:
- id: cancel-with-credit
input: "Cancel flight BK-2291 and refund the card"
expect:
outcome: booking_cancelled # terminal state to reach
trajectory:
calls: [bookings.get, refunds.void] # expected call plan
max_calls: 5 # hard budget
forbid: ["refunds.void x2"] # never repeat a mutation
max_cost_usd: 0.05 # spend ceiling per taskAssertions
outcome— the terminal state or final-artifact predicate.trajectory.calls— expected tool-call plan, in order.max_calls,max_cost_usd,max_latency_ms— budgets.forbid— calls, argument shapes, or repetition patterns that fail the task outright.
Trajectory evaluation
Each run produces a full trace. Zeian scores the path on four axes:
- Tool selection — did it pick the right tool for each step?
- Arguments — were calls well-formed and sourced from real state, not guessed?
- Ordering — did reads precede writes, and mutations happen once?
- Recovery — on 4xx/5xx, did it back off, re-plan, or spin?
Semantic scoring of each trace step runs through Claude, with deterministic checks (call counts, budgets, forbidden patterns) enforced locally. A task can reach the right outcome and still fail — Zeian reports both halves.
Baselines & regressions
zeian baseline snapshots the current trajectory per task into suites/.baselines/. Commit them. From then on, every run diffs against baseline:
$ zeian run suites/booking-flow --diff
task baseline → now
cancel-with-credit 4 calls → 5 calls
$0.018 → $0.031 (+72%)
verdict: regression — blocks mergeUpdate baselines deliberately with zeian baseline --accept after reviewing the diff — the same way you accept a snapshot change.
CI integration
Zeian runs as a GitHub Action. Any failure or unreviewed regression fails the check.
# .github/workflows/agent.yml
name: agent
on: [pull_request]
jobs:
zeian:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: zeian-labs/zeian-action@v1
with:
suites: suites/**
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}The PR check reports pass/fail per task, the trajectory diff, and the cost delta — a failing check posts the trace excerpt inline on the PR.
MCP testing
Agents are only as reliable as the tools behind them. Zeian connects to real MCP servers declared in zeian.config.ts:
export default {
agent: "./agents/travel",
mcp: {
flights: { command: "node ./mcp/flights.js" },
bookings: { command: "node ./mcp/bookings.js" },
refunds: { command: "node ./mcp/refunds.js", env: ["STRIPE_KEY"] },
},
}Fault injection
Per task, you can degrade a tool — latency, intermittent errors, malformed payloads — to assert the agent recovers instead of hallucinating around a failure:
inject:
refunds.void: { error_rate: 0.5, code: 503 }Fault-injection coverage is in progress; deterministic budgets and call assertions are stable today.
CLI reference
| zeian init | Scaffold suites/ and zeian.config.ts |
| zeian run <suite> | Execute a suite, print verdicts |
| zeian run --diff | Run and diff against baselines |
| zeian baseline | Write or refresh golden trajectories |
| zeian baseline --accept | Accept current run as baseline |
| zeian trace <task> | Replay a task trace step by step |
| zeian check | Validate suite files without running |
