Harness drift
A skill written for one agent runtime behaves differently in the next. Tool names, output formats, system prompts — everything shifts.
Early access
A skill that works in one AI coding harness doesn't always work in another — and model updates change behavior without warning. Write your evals once, run them across the harnesses you use, and see exactly what broke.
Launch updates only. No spam.
A skill written for one agent runtime behaves differently in the next. Tool names, output formats, system prompts — everything shifts.
The model behind your harness updates, and yesterday's reliable skill starts failing. Silently, with no alert and no changelog you can act on.
Without evals you find out from your users. With evals tied to a single harness, you only ever see half the picture.
Write evals in code next to your skills: tasks, assertions, graders. Plain files, versioned in your repo.
Execute the same suite across each harness and model you target — on your machine or in CI.
Compare results side by side. See which harness, which model, and which skill regressed — down to the failing case.
Fix the skill, re-run, watch the delta. Ship with proof it works everywhere you care about.
Keep prompts and skills reliable as harnesses and models move under your feet. Know the moment a change breaks something.
Gate skill changes in CI the way you gate code. Review a scorecard before anything reaches production users.
We're building the eval layer for agent skills. Early access goes to the waitlist.
Request access