Evidence: vendor_claim
Anthropic: Automated command review in a coding agent against manual human approval
Across a study of 1,053 testers, humans manually reviewing commands blocked 13.6% of dangerous ones while the automated classifier blocked 89%. Teams and Enterprise users on auto mode are reported to ship roughly 25% more pull requests, with Adobe, Nuro, Gusto and Garner Health named as customers using it.
The strongest-methodology claim of the week and still vendor-run on a product the vendor was making default the same day, which is exactly the conflict to hold in mind. The sample size and disclosed method put it well above the usual vendor assertion, and the direction is consistent with everything known about alarm fatigue in other domains. Use it as the justification for measuring your own approval-catch rate — a cheap exercise almost nobody has done — rather than as a number to cite in a risk register. The 25% pull-request figure is vendor analytics with no baseline disclosed and should not be carried into a business case.
Evidence: vendor_claim
OpenAI and Cerebras: Security investigation loop under low-latency frontier inference
A security investigation workflow reported to drop from one to two hours down to ten to fifteen minutes under the Ultrafast preview, with customer quotes from Jane Street and Podium; throughput cited at up to roughly 14x and around 750 tokens per second.
Compelling and unverifiable in equal measure. No methodology is disclosed, the baseline is anecdotal, and the preview has no published price or availability date, which means even a true result cannot yet be planned against. The useful signal is directional: if latency collapses by an order of magnitude, interactive agent loops become viable in workflows that currently batch, and that changes the user experience question before it changes the cost question. Wait for pricing before redesigning anything.
Evidence: benchmark
FlowScout (academic): Synthesizing reusable agent workflows from historical execution logs
Mining execution logs into a process graph and refining it against execution feedback with Monte Carlo tree search improves tool correctness and execution quality against named baselines including PM4Py, ReAct and AFlow.
Independently measured against named baselines with the method disclosed, which puts it above every vendor claim this week on evidence quality and below all of them on deployment realism — these are academic benchmarks, not production workloads. The transferable idea is cheap to test: if you have historical execution logs for a workflow your agents keep improvising, mine them into a frozen path and measure whether deviation rate drops. That is a one-sprint experiment with a clear success criterion.
Evidence: benchmark
MAP-Graph (academic): Provenance-aware shared memory across multiple cooperating agents
Applying permission, path-trust and action gating as runtime properties of shared agent memory improves task success and accuracy across a 2,700-task suite, with ablations isolating each control's contribution.
Measured with ablations, which is the right way to make this kind of claim, and entirely synthetic, which is the limit on how far to carry it. The finding that matters for practitioners is architectural rather than numerical: if multiple agents share memory, trust needs to be a runtime gate on what one agent may act on from another, not an audit log reviewed afterwards. Teams building multi-agent systems should treat provenance as a design requirement now, because retrofitting it after agents are already writing to shared memory is substantially harder.
Evidence: hype_signal
SpaceXAI: Computer-use agents on internal sales, operations and engineering tasks
Internal workflows including sales outbound, operations invoice handling and engineering bug reproduction cited alongside customer quotes referencing two to three times efficiency gains.
No baseline, no method, no measurement — efficiency multiples quoted in a launch post are marketing, and this publication scores them as such regardless of how plausible the underlying capability is. The capability may well be real and important; the numbers attached to it carry no information. If you pilot this, define your own baseline before the pilot starts, because the vendor has given you nothing to compare against and a post-hoc estimate will be indistinguishable from the launch-post figure.