01
Real-time proprietary sepsis alerts
A deployed alert can add workload while missing most cases.
Negative results
A useful evidence base includes systems that missed cases, created harm, failed to improve the target metric, or worked only inside a bounded capability frontier. These records make the limits visible instead of selecting only favorable outcomes.
Inclusion criteria
Includes cases explicitly identified as negative, failed, or biased, plus cases whose published outcome reports no significant improvement, lower correctness, missed cases, harm, or a material adverse result.
01
A deployed alert can add workload while missing most cases.
02
Fraud-risk triage can cause institutional harm without fair features, recovery rules, or challenge rights.
03
A predictable proxy can still allocate scarce care unjustly.
Current · updated
1 independent, 1 peer reviewed; publication outcomes are verified.
Canonical case →04
A live emergency-call alert may not improve dispatcher recognition despite higher standalone model sensitivity.
05
That access to a general model improves all apparently similar knowledge tasks.
06
Operators set plasma objectives while a simulator-trained policy coordinates all coils on physical hardware.
07
That a sprayer cannot tell a crop plant from a weed while moving through a field, so the only reliable way to control weeds is to broadcast herbicide across the entire field regardless of where weeds actually are.
Current · updated
1 independent, 1 primary; publication outcomes are verified and reported.
Canonical case →08
That preparing a mortgage or HELOC file for a credit decision requires a human to read the document package end to end - classifying documents, extracting figures, calculating income, checking consents and writing the summary - because borrower documents are too variable for a rules engine to handle reliably.
Current · updated
1 independent, 1 primary; publication outcomes are verified and reported.
Canonical case →09
Ambient tools can draft notes in practice, but vendor-specific efficiency and error results require comparison.
10
A blind or low-vision user can get immediate, private scene descriptions before choosing human assistance.
Watch status: verify the cited source and deployment condition before reusing this case.
Watch · updated
Watch status: verify the cited source and deployment condition before reusing this case.
1 peer reviewed; publication outcomes are verified.
Canonical case →11
Border inspectors do not have to rely primarily on random sampling for general-risk imports; historical violations, product attributes, importer data, and international food-safety alerts can target inspection capacity toward higher-risk batches.
Current · updated
2 peer reviewed, 2 primary, 1 independent; scoped inspection-yield outcomes are verified.
Canonical case →