An LLM eval harness you can run in one week
Stop shipping prompt changes on vibes. Harbor’s Eval Harness is a 1-week fixed package ($2,500 sticker · 50/50) that leaves you with a repeatable offline runner and a way to add cases over time.
What “done” looks like
- Eval dataset skeleton + scoring hooks
- Runner for offline regression against a model/prompt pair
- Report format (pass rates, failures, diffs)
- Acceptance: harness runs clean on sample set
- Handoff: how to add cases and gate CI
Explicitly out of scope:
- Human labeling workforce
- Live production traffic shadowing at scale
- Model fine-tuning jobs
- Compliance certification artifacts
Prices and scope above are live from the catalog. Rush sticker is $2,875 when capacity allows. As a Feature Module (P2) add-on, Eval Harness is $2,000 with a shared kickoff — see the packages page note.
One-week shape
- Kickoff: sample prompts / expected outputs (even rough), model endpoint or stub, pass/fail thresholds you care about.
- Build: dataset skeleton, scoring hooks, offline runner, report format with failures and diffs.
- Accept: harness runs clean on the sample set; handoff covers adding cases and gating CI.
Deposit clears before build. Written acceptance against the checklist — see How we work and the AI definition of done template.
Mapped to Gauge
Gauge Eval Harness is the Harbor portfolio case for this package — an illustrative codename study. Metrics there describe timeline, scope, and handoff (for example “1 week build”, “regression runner”, “CI-ready handoff”) — not client logos, attributed testimonials, or model quality claims. Composite studio quotes elsewhere on the site are likewise illustrative.
Starter you can push to GitHub
A minimal local starter (blank golden-set templates + tiny runner stub) lives at
harbor-github-starters/eval-harness-starter/. Org github.com/harborartificial may not exist yet —
Rex creates the org/repo and pushes. The starter README links this guide and the live Eval Harness package card.
Next step
Bring a rough golden set and your model stub. Email [email protected] or open the package card.
Harbor Artificial is not Harbor Cloud, tryharbor.ai, or Protected Harbor. We refuse HIPAA/clinical work, regulated finance systems of record, realtime voice, and open-ended founding-engineer retainers. No profit guarantees. No invented eval metrics.
FAQ
Guide FAQ
How long does an Eval Harness take?
Eval Harness is a one-week fixed package with sticker pricing, written acceptance, and a handoff on how to add cases and gate CI. Rush may be available when capacity allows.
What does Harbor ship in the harness?
A dataset skeleton and scoring hooks, an offline runner against a model or prompt pair, a report format (pass rates, failures, diffs), and acceptance that the harness runs clean on a sample set. Not a human labeling workforce or compliance certification.
How does this relate to Gauge?
Gauge is the Harbor portfolio case study for the Eval Harness package — an illustrative codename study showing timeline, scope, and handoff style. Not an attributed client logo or performance claim.
Want to see if a package fits?
Share what you’re trying to ship, your timing, and any constraints you already know. We’ll reply from [email protected].