AI writes the memo in nine seconds.
Deciding whether it's right is still your job.

Read the artifact below and make the call. Judgment Gym scores your reasoning and shows you what most people miss.

  1. 01

    Read the artifact

    A launch memo, a roadmap tradeoff, an AI analyst's pricing recommendation. Written to look finished, with the problem buried where a confident reader skims past it.

  2. 02

    Make the call

    Commit to a decision and defend it. There is no correct option — a different defensible call, argued well, scores just as highly. The reasoning is what gets read.

  3. 03

    See what you missed

    A score on five dimensions with the evidence quoted back at you, then the thing most people miss on this artifact. That last part is the rep.

critiqueDifficulty 2/3~12 min

The launch memo that reads perfectly

The artifact

From: Priya Raman, Product To: leadership@ Subject: Launch recommendation — Scheduled Exports (GA, Nov 4) Attachments: none

(Drafted with Claude from my notes, then edited. Flagging in case anyone wants the raw version.)

Recommendation

We should take Scheduled Exports to GA on November 4. The feature has been in beta since September, the engineering work is complete, and the market window favours moving before the Q4 procurement cycle closes.

Background

Scheduled Exports lets customers configure recurring CSV and Parquet deliveries to S3, GCS, or an SFTP endpoint on an hourly, daily, or weekly cadence. It addresses the most requested item in our enterprise feedback channel over the last three quarters.

Market context

Data portability has become table stakes in our category. Every serious competitor now ships some form of scheduled delivery, and buyers increasingly treat its absence as a disqualifier in RFPs. Analysts expect the data-integration segment to continue growing at a healthy clip through 2027 as enterprises consolidate their reporting stacks. Our win rate against competitors who offer this is materially lower than against those who don't.

Beta results

  • 31 accounts enrolled, 24 configured at least one export
  • 4,100 successful export runs, 22 failures (99.5% success rate)
  • Median setup time: 6 minutes
  • Support tickets: 9, all resolved, none escalated
  • Feedback has been strongly positive. Several beta customers told us this was "exactly what we'd been waiting for," and two said they'd expand usage once it's generally available.

Engineering readiness

Backend is complete and has been running in production for the beta cohort for seven weeks. Load testing shows headroom to roughly 8× current export volume. The remaining work is a settings-page polish pass (2 days) and documentation (3 days), both in flight. Ben's team has confirmed the timeline.

Pricing

Included in the Growth and Enterprise tiers at no additional charge. We considered metering by export volume but concluded that adding a meter would complicate the sales conversation at a moment when we want the feature to be a reason to buy rather than a reason to negotiate.

Risks

RiskMitigation
Documentation slips past Nov 4Docs team has committed; Priya to track daily
Export volume spikes at GALoad tested to 8×; autoscaling configured
Support load increasesRunbook drafted; support briefed Oct 28
SFTP endpoint variability across customersBeta covered 6 SFTP configurations

Success metrics

We'll measure adoption at 30 and 90 days, targeting 15% of eligible accounts configuring at least one export by day 90. We'll also track the win-rate improvement in deals where scheduled export appears as a requirement.

Ask

Approval to proceed with the November 4 GA date and the announcement plan attached separately.


Your task

You are on the leadership list receiving this memo. Write a critique: identify the most serious problems with it as a decision document, and say what you would need before approving. Rank what you find — not everything you notice matters equally. Aim for 250–450 words.

150 more characters before you can submit

What comes back

The score is the least useful part

Every response is scored on the same five dimensions, so your reps stay comparable. The weights shift by domain — reading a claim isn't the same skill as choosing between two roadmap bets — but the dimensions never change. Open one and you get the evidence, not a grade.

Example, not a live result. A written-out response to Two reviewers, three times the PRs, scored the way yours would be. The answer key below is this scenario's real one — so if you want to try that one cold, skip the last two sections.

The call you made

Buy Sentinel and route review through it.

You were 3/5 confident.

31% of people made the same call.

We don't score the call — there is no correct option, and a different defensible call argued well scores just as highly. The number below is your reasoning.

Weighted score · Two reviewers, three times the PRs

3.85/ 5

Provisional

These scores are provisional. The evaluator is an AI model, and we haven't yet finished checking its scores against human ones. Treat the reasoning below as the useful part and the number as a rough signal.

What worked

You priced Sentinel's 62% against what review actually is today rather than against an idealised careful human — Tomas approving 60% of his queue because the AI left no comment. That comparison is the judgment this scenario is built around, and most responses argue the number in the abstract instead. You also refused to let "humans review the material tier" stand as an oversight design without a sampling rate and an escalation path.

The one change that would raise this most

Name what your call costs. You chose Sentinel and defended it well, but never said what you are accepting in exchange — an undetected-defect rate you cannot yet measure, and a vendor benchmark you cannot audit. A decision that reads as free reads as unconsidered.

By dimension

Clarityweight 0.154Strong

Nadia gets a decision in the first sentence and two conditions attached to it, which is what she asked for and what she can defend to Marcus on the 14th.

Buy Sentinel. Priya's objection is arithmetically right and strategically wrong.

Do this
Say what happens on the 14th. She has to tell Marcus something specific, and you left her to compose it.

Tradeoff awarenessweight 0.252Gestured at

The costs of the two rejected options are named, but not the cost of the one you chose. Picking an option and treating it as free is not making the decision.

Do this
Write the sentence you would least like quoted back to you in six months, then keep it.

User empathyweight 0.103Competent

Tomas and Priya appear as evidence for your argument rather than as two people who are still reviewing thirty PRs a week after your decision lands. Nadia escalated their workload; your call does not change it.

Tomas approving 60% of his queue because the AI left no comment.

Do this
Say what their week looks like on the other side of this. If it looks the same, say that too.

Strategic reasoningweight 0.154Strong

You caught that the dashboard cannot see the failure Nadia is worried about, and made measurement a condition rather than a follow-up. That turns an irreversible bet into a reversible one.

Instrument defect escape rate split by author, because you said you don't know if we're fine.

Do this
Say when you would revisit. A measurement with no review date becomes a dashboard nobody opens.

AI critique qualityweight 0.355Expert

You read the 62% as a comparison rather than as a number, and identified what it is actually being compared against — a review process that stopped existing in March. You also treated the vendor's benchmark as a seller's claim, and refused the material-tier hand-wave without a sampling rate, an escalation trigger and a detection path.

62% isn't competing with a careful human, it's competing with Tomas approving 60% of his queue.

Do this
Nothing load-bearing is missing here. The next increment is naming who audits the audit — the sample you take to check Sentinel itself.

What most people miss— read this one

The common competent answer spots that 62% recall is weak, notices that Sentinel is self-interested, and lands on Option 3 or on some blend — "adopt Sentinel but keep human review on critical paths." That's a defensible call and it will score around a 3. It is also the answer the artifact is built to make attractive: Option 3 is the one the sympathetic character wants, and rejecting a vendor's own benchmark feels like rigour.

What it misses: the comparison isn't 62% versus a careful human. It's 62% versus Tomas approving 60% of his queue because there was no comment. Priya's objection is arithmetically correct and strategically wrong, because she is pricing Sentinel against a review process that stopped existing sometime in March. The artifact says so in one line, in the middle of a paragraph about something else. Responses that read the 62% as a number rather than as a comparison have skipped the judgment entirely — and this is what separates a 3 from a 5 on ai_critique_quality.

The second thing most miss: "humans review the material tier" is the same failure mode as "a human reviews it," which is what everyone in this artifact already agrees is broken. It gets a pass because it arrives inside the expensive option rather than as an excuse. An oversight claim is unfinished until it names the sampling rate, what the reviewer looks for, the escalation trigger, and the after-the-fact detection path. Nobody in the memo has named any of the four, including the vendor.

Third, and most transferable: the highest-value action in this memo is not any of the three options. Nadia writes "I genuinely do not know whether we are fine" and then presents three expensive bets that all depend on the answer. The instrumentation that would tell her — defect escape rate split by author, time-to-approve on zero-comment approvals — costs almost nothing and makes every other choice reversible. Strong responses notice that a decision made blind should buy sight first, and that this doesn't conflict with also making the call.

The plausible wrong answer is to reject Option 1 on the grounds that the vendor is self-interested. The vendor is self-interested, and reading claims against the seller's position is a real heuristic — but here it points at their tiering (unpublished, unmeasured), not at the 62%, which is the number that most damages their own "vanishingly little human participation" pitch. A seller publishing a figure that undercuts their own marketing is the part of their evidence you should trust most. Responses that use "they're biased" as a conclusion rather than as a place to look score poorly on ai_critique_quality.

Fourth, quietly: almost nobody prices Option 3's hidden cost. Funding verification means pulling Tomas — the reviewer who hasn't started rubber-stamping — off review for a quarter. The safe-sounding option degrades oversight faster than the risky one, and the artifact never says so.

An exemplar response— one good answer, not the only one

Nadia — taking Option 1, with two conditions, and I'll tell you why the framing is slightly off.

The reason it's not close. Option 2 is already what's happening, badly. Capping merges at 60 doesn't restore review depth; it just moves the rationing from Priya's attention to the backlog, and we lose the capacity as well. And Option 3 is right on a two-year horizon and useless on a two-week one — it also takes Tomas, our more careful reviewer, off review for a quarter, which makes the thing we're worried about worse while we fix it.

But Priya is right about the 62%, and Raj is right too. They're arguing about different denominators. 62% is not "how safe are we" — it's "how much does the tool catch of what it was tested on, on a benchmark the vendor chose." Two in five serious bugs walking would be alarming if the alternative were an attentive human. The actual alternative is Tomas approving 60% of his queue because there was no comment. That's not review; we already have automated approval, we just haven't admitted it or measured it. Sentinel is not a decrease in rigour. It's the same rigour, made explicit and cheaper.

What I'm not buying is their oversight design. "Humans on the material tier" is a hope, not a mechanism, until it says four things: what fraction of PRs land in it, what the reviewer is looking for, what triggers escalation, and how we catch what slipped after the fact. Ask Raj to get that in writing before we sign. And their tiering is the part they have not published numbers on — the 62% tells us nothing about whether the tier assignment is any good.

Two conditions.

  1. Sentinel must not be running the model that wrote the code. Their own number says cross-model beats same-model by 8 points, and the mechanism is that a model misses the bugs it makes. If the config can't guarantee that, the deal doesn't work.
  2. Priya's remaining human review gets aimed, not sampled. Not 10% at random — the classes where this fails: anything touching money movement, anything that changes an auth boundary, anything where the diff is small but the behaviour change isn't. Her queue should get smaller and more deliberate, not just smaller.

And the thing you're not seeing. You said you can't tell whether quality dropped. That's the actual emergency, and it's cheaper to fix than any of the three options. Start splitting defect escape rate by agent-authored vs human-authored now, and track time-to-approve on zero-comment approvals — that second number is the only one that would have shown us Tomas's problem before he mentioned it. Do that this week regardless of the Sentinel decision. If we're wrong about Option 1, that's the instrument that tells us, and without it we're committing $49k a year on a benchmark a vendor published about itself.

For the 14th: we're making review explicit and instrumented rather than implicit and unmeasured, and we'll know within a quarter whether it held.

Applies: MM-WA-07 (verification cost sets the delegation boundary), FP-WA-02 (rubber-stamp review — already happening here), FP-WA-01 (ungated human in the loop — Sentinel's "material tier"), H-WA-04 (an oversight design names sampling rate, target, escalation, and after-the-fact detection), H-WA-12 (sample by failure class, not at random), H-WA-18 (read capability claims against the seller's position), MM-WA-03 (humans monitor reliable processes badly).

The library

12 scenarios, all scored, no signup

See all 12

One scenario a week, in your inbox

A new artifact to tear apart, plus what most people missed on the last one. No cohort, no curriculum — reps.

One email a week. Unsubscribe any time.