Documentation

Zeian docs.

Zeian is in alpha. The surface described here is implemented locally and evolving quickly — where a feature is still in progress, the page says so.

Quickstart

Install

$ npm install -g @zeian/cli
$ zeian --version
zeian 0.4.2 (alpha)

Initialize a suite

Run zeian init inside the repository that holds your agent. It creates suites/ with a starter task and a zeian.config.ts pointing at your agent entrypoint.

$ cd my-agent && zeian init
created suites/smoke.yaml
created zeian.config.ts

Run

$ zeian run suites/smoke.yaml
✓ greets-and-routes      pass · 3 calls · $0.008
1/1 pass

A green run means every task reached its expected outcome and its trajectory stayed inside the declared budget.

Concepts

Suites and tasks

A suite is a YAML file describing a family of realistic tasks. Each task declares an input, the tools the agent may use, and what success means. Suites live in your repo — they are code, reviewed like code.

Trajectory

The ordered record of everything the agent did: tool calls with arguments, retries, model turns, cost, and latency. Zeian asserts on the trajectory, not just the final answer — an agent that calls a refund twice and guesses its way to the right output is a production incident, not a pass.

Verdict

Per-task result: pass, fail, or regression. A regression means the task passed but its behavior drifted from the stored baseline — more calls, a different tool order, higher cost.

Baseline

A committed golden trajectory per task. Baselines are what make diff review possible: every run is compared against them and the drift is shown in the report.

Suite format

An annotated suite:

suite: booking-flow            # suite id
agent: ./agents/travel         # agent entrypoint
tools: [flights, bookings, refunds]  # MCP servers allowed

tasks:
  - id: cancel-with-credit
    input: "Cancel flight BK-2291 and refund the card"
    expect:
      outcome: booking_cancelled    # terminal state to reach
      trajectory:
        calls: [bookings.get, refunds.void]  # expected call plan
        max_calls: 5                 # hard budget
        forbid: ["refunds.void x2"]  # never repeat a mutation
        max_cost_usd: 0.05           # spend ceiling per task

Assertions

  • outcome — the terminal state or final-artifact predicate.
  • trajectory.calls — expected tool-call plan, in order.
  • max_calls, max_cost_usd, max_latency_ms — budgets.
  • forbid — calls, argument shapes, or repetition patterns that fail the task outright.

Trajectory evaluation

Each run produces a full trace. Zeian scores the path on four axes:

  • Tool selection — did it pick the right tool for each step?
  • Arguments — were calls well-formed and sourced from real state, not guessed?
  • Ordering — did reads precede writes, and mutations happen once?
  • Recovery — on 4xx/5xx, did it back off, re-plan, or spin?

Semantic scoring of each trace step runs through Claude, with deterministic checks (call counts, budgets, forbidden patterns) enforced locally. A task can reach the right outcome and still fail — Zeian reports both halves.

Baselines & regressions

zeian baseline snapshots the current trajectory per task into suites/.baselines/. Commit them. From then on, every run diffs against baseline:

$ zeian run suites/booking-flow --diff
task                baseline → now
cancel-with-credit  4 calls → 5 calls
                    $0.018  → $0.031  (+72%)
verdict: regression — blocks merge

Update baselines deliberately with zeian baseline --accept after reviewing the diff — the same way you accept a snapshot change.

CI integration

Zeian runs as a GitHub Action. Any failure or unreviewed regression fails the check.

# .github/workflows/agent.yml
name: agent
on: [pull_request]
jobs:
  zeian:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: zeian-labs/zeian-action@v1
        with:
          suites: suites/**
        env:
          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}

The PR check reports pass/fail per task, the trajectory diff, and the cost delta — a failing check posts the trace excerpt inline on the PR.

MCP testing

Agents are only as reliable as the tools behind them. Zeian connects to real MCP servers declared in zeian.config.ts:

export default {
  agent: "./agents/travel",
  mcp: {
    flights:  { command: "node ./mcp/flights.js" },
    bookings: { command: "node ./mcp/bookings.js" },
    refunds:  { command: "node ./mcp/refunds.js", env: ["STRIPE_KEY"] },
  },
}

Fault injection

Per task, you can degrade a tool — latency, intermittent errors, malformed payloads — to assert the agent recovers instead of hallucinating around a failure:

    inject:
      refunds.void: { error_rate: 0.5, code: 503 }

Fault-injection coverage is in progress; deterministic budgets and call assertions are stable today.

CLI reference

zeian initScaffold suites/ and zeian.config.ts
zeian run <suite>Execute a suite, print verdicts
zeian run --diffRun and diff against baselines
zeian baselineWrite or refresh golden trajectories
zeian baseline --acceptAccept current run as baseline
zeian trace <task>Replay a task trace step by step
zeian checkValidate suite files without running