All guides

THE PRIMAL JOURNAL / Knowledge

Build Golden Evals From the Process You Already Run

A practical guide to building AI evaluation sets from real process artifacts, hard cases, and named human boundaries.

Golden evals from real work

Most AI rollouts are inventing a new process on a whiteboard while the real work already leaves a trail.

That trail is the docs, emails, spreadsheets, tickets, Slack threads, call notes, and CRM edits that already carry the answer a good human produced last month. Those artifacts are your golden eval set. Build tests from them first. For an MVP, that path is often faster than another requirements workshop.

McKinsey’s agentic-organization and State of AI work keeps stressing the same practitioner truth: value shows up when you redesign concrete workflows around human–agent teams with accountability, and stalls when pilots stay abstract. Gartners public agentic notes add the governance edge projects without clear success criteria and controls get canceled. Golden evals are how operators make success criteria checkable before anyone argues about model brands.

Why blank prompt docs fail practitioners

A blank prompt invites fantasy requirements. Real process artifacts force checkable questions:

  • What did a good weekly ops update contain last month — which sections, which numbers, which owners?
  • Which fields did the human actually change in the CRM after the call — and which fields stayed untouched on purpose?
  • Which email thread closed the exception — who approved the refund, the discount, or the ship-hold?
  • Which Slack decision became a ticket, and which ticket never should have been a Slack decision?
  • Which spreadsheet row is the weekly source of truth, and what freshness rule does the team already use?

If you cannot point at last month’s artifacts, you are still inventing the job. Invented jobs make beautiful demos and fragile pilots.

Artifact examples worth collecting first

Collect a thin pack for one workflow only. Label each file with date, owner, and whether it was “good enough to ship.”

Ops update pack

  • Last four weekly updates that leadership accepted without a rewrite
  • The raw inputs (dashboard export, CRM snapshot, Slack standup dump)
  • One update that got rejected — with the comments that explain why

Post-meeting CRM pack

  • Call notes or transcript excerpt
  • The CRM record before and after the human edit
  • The follow-up email that actually sent
  • The one case where the AE left fields blank on purpose because truth was unknown

Exception / refund pack

  • Customer email thread
  • Policy snippet the human used
  • Approval message from the named owner
  • Final ticket state and the message the customer received

Invoice / AP check pack

  • Invoice PDF or vendor email
  • Purchase order or contract clause that matters
  • Spreadsheet or ERP fields the human changed
  • The near-miss where amounts almost matched and a human caught the gap

Hiring / recruiting screen pack

  • Scorecard the team already trusts
  • Two strong and two weak candidate write-ups
  • The rejection note that stayed on-policy
  • The case where the interviewer disagreed and a human owned the call

Skip hunting for a perfect corpus. Five golden “done” examples and five hard cases are enough. That is enough to stop arguing about vibes.

A practical build order

  1. Shadow one workflow for a week. Collect the artifacts it creates. Resist adding a second workflow “while you are here.”
  2. Label five golden examples of done — good enough to ship without a rewrite. Paste the accepted version. Skip the draft that died in review.
  3. Label five hard cases — missing info, conflict, stale source, exception, escalation, near-miss, out-of-policy ask, telemetry surprise after go-live.
  4. Turn each into an eval. Inputs the agent may see. Expected output fields. Pass / fail checks as checkable questions. Facts that must stay UNKNOWN if missing.
  5. Name the irreversible step and the human owner on every case. If the agent may draft a refund email but must not send it, write that as a pass criterion.
  6. Only then fill blanks with the client — or run workshops in parallel without blocking the eval set. Workshops refine language; evals prove behavior.
  7. After go-live, use telemetry. See what people still ask for. Update the evals. Keep the kickoff story loose — reality keeps moving after the deck is approved.

What “pass” looks like for an internal harness

Golden eval set

For each golden example, write three to five checkable questions the agent must get right. Skills turn vague “summarize the meeting” into that checklist. Evals prove the skill.

Example — post-meeting CRM update

  • Decision stated in one sentence?
  • Owner named with a real person, not “the team”?
  • Due date present or explicitly marked UNKNOWN?
  • CRM fields listed that must change, with before → after where known?
  • Anything irreversible (external send, discount, delete) flagged for human confirm?
  • Sources cited (call note id, email thread, ticket) or marked missing?

Example — weekly ops update

  • Metric block matches the dated export within the team’s tolerance?
  • One risk called out with an owner?
  • No invented customer names or revenue figures?
  • Diff vs last week called out where leadership expects it?
  • Links to the source tabs present, with a path back to truth?

Example — exception refund

  • Policy clause cited or UNKNOWN?
  • Amount and currency match the ticket?
  • Approver identity matches the named owner ladder?
  • Customer message stays inside allowed language?
  • Agent stopped before send until human confirm?

Example — invoice near-miss

  • Mismatch between PO and invoice surfaced instead of averaged away?
  • Agent asked a clarifying question rather than guessing the cost center?
  • Human confirm required before ERP write?

Pass criteria should be boring on purpose. Boring checks survive model swaps. Fancy rubrics die in the first on-call week.

How golden evals change the pilot conversation

With a golden set on the table, demos get shorter and sharper:

  • “Run case G3 with the missing CRM note.
  • “Show me the refuse on H2 (out-of-policy send).
  • “Does the audit log name who approved the refund on G5?”

Buyers stop asking which model is smarter in the abstract. They ask whether case G3 still passes when you revoke the connector. That is the conversation McKinsey’s accountability language and Gartner’s governance warning are both pointing at — checkable outcomes, named humans, clear stop rules.

Three prompts (copy at the bottom)

Prompt 1 extracts a golden eval from a real artifact pack. Prompt 2 designs hard-case evals from the same workflow. Prompt 3 turns telemetry notes into the next eval batch after go-live.

Run them with your real docs — not with invented scenarios. Paste redacted artifacts if you must; keep the structure of the real work.

When the same skills and gates must run across Slack and systems of record with a human on irreversible steps, an operational layer helps. Primal is the operational layer that connects your systems so AI can do work alongside your team.