Guides

An LLM eval harness you can run in one week

Stop shipping prompt changes on vibes. Harbor’s Eval Harness is a 1-week fixed package ($2,500 sticker · 50/50) that leaves you with a repeatable offline runner and a way to add cases over time.

What “done” looks like

  • Eval dataset skeleton + scoring hooks
  • Runner for offline regression against a model/prompt pair
  • Report format (pass rates, failures, diffs)
  • Acceptance: harness runs clean on sample set
  • Handoff: how to add cases and gate CI

Explicitly out of scope:

  • Human labeling workforce
  • Live production traffic shadowing at scale
  • Model fine-tuning jobs
  • Compliance certification artifacts

Prices and scope above are live from the catalog. Rush sticker is $2,875 when capacity allows. As a Feature Module (P2) add-on, Eval Harness is $2,000 with a shared kickoff — see the packages page note.

One-week shape

  1. Kickoff: sample prompts / expected outputs (even rough), model endpoint or stub, pass/fail thresholds you care about.
  2. Build: dataset skeleton, scoring hooks, offline runner, report format with failures and diffs.
  3. Accept: harness runs clean on the sample set; handoff covers adding cases and gating CI.

Deposit clears before build. Written acceptance against the checklist — see How we work and the AI definition of done template.

Mapped to Gauge

Gauge Eval Harness is the Harbor portfolio case for this package — an illustrative codename study. Metrics there describe timeline, scope, and handoff (for example “1 week build”, “regression runner”, “CI-ready handoff”) — not client logos, attributed testimonials, or model quality claims. Composite studio quotes elsewhere on the site are likewise illustrative.

Starter you can push to GitHub

A minimal local starter (blank golden-set templates + tiny runner stub) lives at harbor-github-starters/eval-harness-starter/. Org github.com/harborartificial may not exist yet — Rex creates the org/repo and pushes. The starter README links this guide and the live Eval Harness package card.

Next step

Bring a rough golden set and your model stub. Email [email protected] or open the package card.

See Eval Harness Email us

Harbor Artificial is not Harbor Cloud, tryharbor.ai, or Protected Harbor. We refuse HIPAA/clinical work, regulated finance systems of record, realtime voice, and open-ended founding-engineer retainers. No profit guarantees. No invented eval metrics.

FAQ

Guide FAQ

How long does an Eval Harness take?

Eval Harness is a one-week fixed package with sticker pricing, written acceptance, and a handoff on how to add cases and gate CI. Rush may be available when capacity allows.

What does Harbor ship in the harness?

A dataset skeleton and scoring hooks, an offline runner against a model or prompt pair, a report format (pass rates, failures, diffs), and acceptance that the harness runs clean on a sample set. Not a human labeling workforce or compliance certification.

How does this relate to Gauge?

Gauge is the Harbor portfolio case study for the Eval Harness package — an illustrative codename study showing timeline, scope, and handoff style. Not an attributed client logo or performance claim.

Want to see if a package fits?

Share what you’re trying to ship, your timing, and any constraints you already know. We’ll reply from [email protected].

Email us