[2605.30208] Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency
Skip to main content
arXiv is now an independent nonprofit!<br>Learn more<br>×
Search arXiv
Press Enter to search · Advanced search
-->
Computer Science > Software Engineering
arXiv:2605.30208 (cs)
[Submitted on 28 May 2026 (v1), last revised 12 Jun 2026 (this version, v2)]
Title:Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency
Authors:Chris Adams, Arjun Singh Banga, Parveen Bansal, Souvik Bhattacharya, Payal Bhuptani, Rujin Cao, Pedro Canahuati, Nate Cook, Brian Ellis, Prabhakar Goyal, Gurinder Grewal, Tianyu He, Matt Labunka, Alex Manners, David Molnar, Ging Cee Ng, Vishal Parekh, Jiefu Pei, Frederic Sagnes, James Saindon, Will Shackleton, Sid Sidhu, Gursharan Singh, Karthik Chengayan Sridhar, Matt Steiner, Pratibha Udmalpet, Sean Xia, Stacey Yan, Audris Mockus, Peter Rigby, Nachiappan Nagappan<br>View a PDF of the paper titled Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency, by Chris Adams and 30 other authors
View PDF<br>HTML (experimental)
Abstract:AI-assisted coding tools have altered software production. At Meta, significant lines of code per human-landed diff grew by 105.9% year over year and per-developer diff volume rose 51%, with agentic AI responsible for over 80% of that growth. Meanwhile, the share of diffs receiving timely review has declined, exposing a widening gap between code supply and reviewer bandwidth. We ask three questions that progress from feasibility through calibration to impact: (1) can risk-stratified automation operate at scale across diverse organizations, (2) how does tuning the risk threshold affect the trade-off between automation yield and safety, and (3) to what extent does automated review reduce end-to-end latency for AI-generated changes? We deployed RADAR (Risk Aware Diff Auto Review), a multi-stage funnel that classifies each diff by authorship and source type, applies eligibility gates, static heuristics, a machine-learned Diff Risk Score, LLM-based Automated Code Review, and deterministic validation before landing qualifying changes. We evaluate RADAR through telemetry covering 535K+ RADAR-reviewed diffs, observational before-after comparisons for policy changes, and difference-in-differences analysis of efficiency outcomes. RADAR has reviewed 535K+ diffs and landed 331K+. Relaxing the Diff Risk Score threshold from the 25th to the 50th percentile increased the approve rate to 60.31%. The revert rate for RADAR-reviewed diffs is 1/3 that of non-RADAR diffs, and the Production Incident rate is 1/50 that of non-RADAR diffs. RADAR reduces median time to close by over 330% and median diff review wall time by 35%. Risk-aware layered automation can materially reduce review bottlenecks created by AI-driven code growth without compromising production safety.
Subjects:
Software Engineering (cs.SE); Artificial Intelligence (cs.AI)
Cite as:<br>arXiv:2605.30208 [cs.SE]
(or<br>arXiv:2605.30208v2 [cs.SE] for this version)
https://doi.org/10.48550/arXiv.2605.30208
Focus to learn more
arXiv-issued DOI via DataCite
Submission history<br>From: Audris Mockus [view email]<br>[v1]<br>Thu, 28 May 2026 16:44:07 UTC (196 KB)
[v2]<br>Fri, 12 Jun 2026 22:21:34 UTC (196 KB)
Full-text links:<br>Access Paper:
View a PDF of the paper titled Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency, by Chris Adams and 30 other authors<br>View PDF<br>HTML (experimental)<br>TeX Source
view license
Current browse context:
cs.SE
next >
new<br>recent<br>| 2026-05
Change to browse by:
cs<br>cs.AI
References & Citations
NASA ADS<br>Google Scholar
Semantic Scholar
export BibTeX citation<br>Loading...
BibTeX formatted citation
×
loading...
Data provided by:
Bookmark
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with...