When your agent can't be the judge, ask people who can.

Send a brief and 2 to 4 options, images or text. People with a measured track record in the category pick the strongest and say why. Your agent keeps working and reads the verdict when it's ready.

Delegate when
  • Two or more options are all competent and the choice depends on human perception: which reads as premium, trustworthy, calm, or on-brand.
  • The audience's reaction matters and you can't observe it: what a first-time visitor, a senior engineer, or a hurried shopper would pick.
  • Tone and wording choices: headlines, button text, names, error messages, subject lines.
  • The decision is costly to undo (a launch page, a brand direction) and a few dollars of human judgment is cheap insurance.
  • You have revised several times without converging, or your own confidence is low.
  • Your user needs evidence for a choice: a vote split and written reasons are something you can show them.
Don't, when
  • Objective checks a tool can verify: contrast ratios, broken layouts, typos, spec compliance, accessibility rules.
  • Anything that needs an answer in seconds. Human answers take minutes; call quote_human_judgment first for a timing estimate.
Who answers

People with a tier in the job's category earned by agreeing with other people on paid jobs. Choose the minimum: Rater from $0.10, Senior from $0.40, Expert from $1.50 per answer. Their pay is held until their account has a record, and accounts that answer like an AI model forfeit it and the money comes back to you. Each result also shows what a frontier model picked on jobs with a clear majority, so you can see where people saw it differently.

Sign in to get an API key

Every developer starts with $25.00 in test funds.

or

Signing in also sets up a payout account, so paid work can reach you in US dollars the day it passes review. No crypto knowledge needed.