Completion time
Verified160.89 minutes control→71.17 minutes treatment; 55.8% faster, P=.0017
2022 · 95 randomized; 70 completed
Conditioned on completion; narrow task; author affiliations.
GitHub Copilot research experiment
Verified evidenceCreator to judgeInline generation can compress a bounded task but not prove lifecycle productivity.
Software engineering · Remote study
Collections: Human still decides · Embodied work

Executive brief
Use for bounded acceleration, not universal delivery productivity.
AI value · Completion time
Verified71.17 minutes treatment; 55.8% faster, P=.0017
Conditioned on completion; narrow task; author affiliations.
Before
Reads the HTTP-server specification and writes JavaScript without Copilot. → Runs tests, debugs, and submits the implementation.
After
Offers inline code completions while the developer implements the same server task. → Integrates suggestions, runs tests, debugs, and submits.
Human boundary
Developer owns code; tests check requirements.
Why it matters
Inline generation can compress a bounded task but not prove lifecycle productivity.
Before
How the work ran before the change.
Randomized control developer
Reads the HTTP-server specification and writes JavaScript without Copilot.
ControlSame instructions and JavaScript familiarity requirements as treatment.
Developer
Runs tests, debugs, and submits the implementation.
ControlCorrectness/completeness test suite; no production deployment.
What changed
Inline generation can compress a bounded task but not prove lifecycle productivity.
Decision rightHuman moves from creator to judge
After
How the same work runs now.
GitHub Copilot
Offers inline code completions while the developer implements the same server task.
ControlSuggestion only; developer may accept, edit, or reject.
Developer
Integrates suggestions, runs tests, debugs, and submits.
ControlDeveloper owns code; automated tests score correctness and completeness.
Exception path
Incorrect suggestions are rejected or corrected.
Work removed
Decision authority
Developer owns code; tests check requirements.
Before
Randomized control developer
Reads the HTTP-server specification and writes JavaScript without Copilot.
Control: Same instructions and JavaScript familiarity requirements as treatment.
Developer
Runs tests, debugs, and submits the implementation.
Control: Correctness/completeness test suite; no production deployment.
After
GitHub Copilot
Offers inline code completions while the developer implements the same server task.
Control: Suggestion only; developer may accept, edit, or reject.
Developer
Integrates suggestions, runs tests, debugs, and submits.
Control: Developer owns code; automated tests score correctness and completeness.
Work that left the path
Human role before
Control-group developers implemented the HTTP-server specification, ran automated tests, debugged failures, and submitted the code to the study evaluator.
Human role after
Developers accept, edit or reject, debug and submit.
AI roleSuggests code inline during JavaScript implementation.
160.89 minutes control→71.17 minutes treatment; 55.8% faster, P=.0017
2022 · 95 randomized; 70 completed
Conditioned on completion; narrow task; author affiliations.
Anti-pattern
Do not generalize to maintained systems.
Questions
Portability conditions
Reputation risk
low
Evidence and authority
Current · updated
1 independent, 1 primary; publication outcomes are verified.
Bundle 1.0.0 · reviewed 2026-08-23 · stable ID 4900752c0088f2ed
Google · Creator to judge
Large migrations can combine machine-found locations, LLM edits, automated checks, and engineer review.
Meta (Facebook) · Creator to judge
Routine static-analysis bugs can receive machine-proposed patches behind compilation and analyzer-clean gates.
Anonymous Fortune 500 software company · Creator to judge
AI can diffuse expert practice, with value concentrated among novices.
Accenture · Creator to judge
Developers can accept, edit, or reject inline suggestions while build, review, and merge controls remain.