Skip to main content

Clarity Harness · Open Source · Launching Soon

Compound your context.

Every agent run makes the next one better. Incubate your agents — verdicts, not rewrites — until they earn autopilot, and your taste compounds right in the flow of work.

1,437 builders on the waitlist · launching soon · free & open source

Manage your agents

Old world vs new world

Agents made you fast.
They didn't make you sure.

Without Clarity

You bought agents. You got slop.

  • Drafts pile up, unreviewed and unowned
  • Your best people babysit every artifact
  • “Done” quietly means “probably”

With Clarity

Every run in one queue, review state attached.

The harness captures every agent run — story drafts, compliance checks, recon reports — grouped by agent with its review state on the row. Nothing ships as a silent maybe.

The harness runs queue — agent runs for a payments platform with pending-review states
The runs queue — every agent, every run, every review state.

Without Clarity

You’re the eval. Forever.

  • Seniors read everything, line by line
  • Or rubber-stamp at 5pm under deadline
  • Judgment is the bottleneck — it doesn’t scale

With Clarity

Verdicts in minutes, receipts attached.

Approve, accept with edit, reject, defer — one keystroke each. The chain, the diff, and the source sit one tab away, so judging a run takes minutes, not meetings.

The harness judgment view — verdict queue and one-keystroke decisions on a chargeback story
The judgment view — verdict queue, receipts, one-keystroke decisions.

Without Clarity

Your corrections evaporate.

  • The same mistake, caught every week
  • Fixes die in Slack threads
  • Agents never remember being wrong

With Clarity

Corrections carried forward. Forever.

Every correction is a reviewable diff — title, summary, acceptance criteria — banked into a versioned dataset that feeds the next regeneration. Fix it once and the whole harness learns.

The harness delta view — an artifact version diff with additions highlighted, ready to approve or annotate
The delta — every version, every change, one diff away.

Without Clarity

You review forever.

  • Trust never accumulates
  • Week 12 reviews look exactly like week 1
  • Autonomy stays a demo

With Clarity

Agents level up — all the way to autopilot.

Advisor → Copilot → Autopilot. Promotion needs the full checklist, not a headcount: gates cleared with annotated traces, golden sets, and review coverage — the harness tracks the XP.

Agent standing page — Advisor, Copilot, and Autopilot tiers with promotion gates, XP, and the error-analysis quest log
Agent standing — tiers, promotion gates, and the quest log to autopilot.

Without Clarity

Your prompts rot in a doc.

  • The “latest” system prompt lives in five places
  • Nobody knows what changed, or why
  • Improvements can’t be rolled back — or repeated

With Clarity

The harness itself is versioned.

System prompt, skills, acceptance-criteria templates — every regeneration ships as a new harness version with a reviewable diff. v0.1 at 21 traces, v0.3 at 87: you can see exactly how your agent got good.

The PM Story Writer harness page — versions over time with trace counts, and the live system prompt
The Harness tab — versions over time, the live system prompt, the grading rubric.

Speed without compounding is just expensive slop.

The promise

Every verdict compounds.

The harness closes the loop: capture the run, deploy the agent, review with a verdict, learn from the correction. Today's annotation is tomorrow's default. Week one feels like review. Week six feels like leverage.

Clarity Loop

Capture

Every run recorded, receipts attached

Deploy

The agent runs your workflow

Review

You give the verdict in minutes

Learn

Corrections banked into the next run

Continuously improving

Proof

We've helped real businesses with AI — from YC startups to $2B companies.

We can help you too.

What ships at launch

Everything we run internally. All of it. Free.

harness_ui

The full harness UI

Mission control, run views, judgment queue — the exact interface in the screenshots above.

run_capture

Runs with receipts

Every run recorded: source input, agent chain, tool calls, decisions taken and not taken.

verdicts

The verdict system

Approve, edit, reject, defer — annotations persist as structured data, not Slack threads.

datasets

Versioned datasets

Corrections become a dataset with lineage — dev, staging, prod — that regenerates your agents better.

mcp_serve

MCP serve

Point Claude Code or any coding agent at your harness — orchestrator calls the agent, agent calls the skill.

setup_guide

Our setup playbook

The internal guide we used to stand up our own harness. Running the same afternoon.

Built in the open

A community for builders who believe agents deserve better than vibes.

The harness ships free and open source because the thesis is bigger than a tool. This is the Clarity bet — and we're building a community of AI builders who share it:

Subjectivity is the missing primitive

Autonomy doesn’t come from bigger models. It comes from knowing whose goals, whose judgment, and whose context a run answers to.

Alignment is layered

Individual, team, and organization goals aren’t the same thing — a real harness reconciles all three instead of pretending one prompt covers it.

Closed loops beat bigger prompts

True autonomous agents are grown, not prompted: run → verdict → learn → run again, with the human judgment banked every cycle.

If that's your thesis too, don't just star the repo — come build the loop with us.

Scaling past your own harness? Company-wide harnesses are the layer we build with design partners — see the company brain →

Open source · launching soon

Founding builders get in first.

The waitlist gets the repo before it's public, our internal setup playbook, and founding-builder credit in the repo. Free, open source, no license keys.

1,437 builders on the waitlist · launching soon