-
arXiv:2605.20926
"MemConflict," arXiv:2605.20926.
Cited in 4 locations
-
DOI:10.1145/3241743
2022](https://doi.org/10.1109/TSE.2022.3174092)), the ABC framework for software-engineering research (Stol and Fitzgerald 2018), and work on construct validity in software engineering (Sjøberg and Bergersen 2023).
Cited in 1 location
-
DOI:10.1109/TSE.2022.3176725
2022](https://doi.org/10.1109/TSE.2022.3174092)), the ABC framework for software-engineering research (Stol and Fitzgerald 2018), and work on construct validity in software engineering (Sjøberg and Bergersen 2023).
Cited in 1 location
-
arXiv:2606.14594
A comparative policy analysis by Manita and Amari (2026) examined six open-source organizations and used documented agent incidents to derive this six-dimension taxonomy.
Cited in 1 location
-
arXiv:2403.04811
A controlled analysis by Riddell and colleagues (2024) measured overlap between open corpora and HumanEval or MBPP, then found a performance difference between contaminated and clean subsets.
Cited in 1 location
-
arXiv:2403.19114
A controlled coding comparison by Xia and colleagues (2024) found that LLM-evolved HumanEval variants reduced scores by ~39% on average and reordered leaderboard positions.
Cited in 1 location
-
arXiv:2505.13353
A controlled counterfactual study across 10 models by Štorek and colleagues (2025) found near-perfect position-independent lexical recall while semantic recall deteriorated for centrally positioned code.
Cited in 1 location
-
arXiv:2402.02823
A controlled evasion study by Dekoninck and colleagues (2024) then demonstrated that Evasive Augmentation Learning could raise tested scores while escaping every evaluated detector.
Cited in 1 location
-
arXiv:2502.05167
A controlled long-context evaluation by Modarressi and colleagues (2025) removed literal matching cues and found that 10 of 12 models claiming >=128K windows fell below 50% of their short-context baseline at 32K tokens.
Cited in 1 location
-
arXiv:2402.01781
A controlled multiple-choice comparison by Alzahrani and colleagues (2024) found ranking movements of up to 8 positions under answer-choice reordering and alternative scoring methods.
Cited in 1 location
-
arXiv:2506.12286
A controlled study by Liang and colleagues (2025) found up to 76% file-path identification from SWE-Bench-Verified issue text alone, with much lower accuracy on held-out repositories.
Cited in 1 location
-
arXiv:2310.17623
A controlled study by Oren and colleagues (2023) showed that comparing canonical and shuffled dataset orders through log probabilities can provide statistical contamination evidence, including for datasets seen only a few times during pretraining.
Cited in 1 location
-
arXiv:2603.19328
A controlled tau-bench evaluation by Sah and colleagues (2026) found interception as high as 94%, while strictly safe goal attainment remained below 5% in most settings.
Cited in 1 location
-
arXiv:2208.13068
A database-integrated function framework studied by Kraft and colleagues (2022) provides directional evidence for using the database as the execution runtime so the log and effects share a commit boundary.
Cited in 1 location
-
arXiv:2602.23905
A descriptive observational study by Ammar Asdaque and colleagues (2026) examined 22,953 pull requests from 1,719 vibe coders.
Cited in 1 location
-
arXiv:2605.26563
A directional author-run evaluation by Wang and colleagues (2026) found that diagnosis effectiveness degraded as repository-level coding trajectories grew longer and noisier.
Cited in 1 location
-
arXiv:2512.22087
A directional benchmark result from Liu and colleagues (2025) found that agent-invoked structured compression under a bounded context budget outperformed static compression and ReAct-style baselines on SWE-bench Verified.
Cited in 1 location
-
arXiv:2407.07565
A directional benchmark-construction study by Matton and colleagues (2024) points toward anti-searchability review as a defense spanning these channels.
Cited in 1 location
-
arXiv:2511.00197
A directional observational study across OpenHands, SWE-agent, and Prometheus on SWE-bench found that failed trajectories were consistently longer and more variable than successful ones, even while file localization remained sound, as reported by Majgaonkar and colleagues (2025).
Cited in 1 location
-
arXiv:2604.09963
A directional preprint by Bindschaedler (2026) reports that typed plans validated by a microkernel reduced agent-caused harm in simulation and online evaluation over industrial traces with fault injection.
Cited in 1 location
-
arXiv:2503.12374
A directional process-mining study by Chen and colleagues (2025) examined 3,977 solving trajectories and 3,931 test logs from 8 agents on 500 SWE-bench issues.
Cited in 1 location
-
arXiv:2411.13323
A directional study by Ramos and colleagues (2024) found memorization signals concentrated in particular pairings, including codegen-multi on Defects4J, while newer models trained on larger corpora, such as LLaMa 3.1, showed weaker signals.
Cited in 1 location
-
arXiv:2607.06184
A directional study by Shu and colleagues (2026) applied this approach to 2,500 trajectories from five production settings on SWE-bench Verified.
Cited in 1 location
-
arXiv:2411.03923
A directional study by Singh and colleagues (2024) suggests that benefit-anchored measures could make competing contamination definitions more comparable and could filter out overlap without a detectable score effect.
Cited in 1 location
-
arXiv:2605.16278
A framework paper by Gaube and colleagues (2026) proposes a standard architecture for effective human oversight and applies it illustratively across domains.
Cited in 1 location
-
arXiv:2510.00328
A grey-literature review by Fawzy and colleagues (2025) examined 518 firsthand accounts from 101 practitioner sources.
Cited in 1 location
-
arXiv:2504.20879
A large forensic observational analysis by Singh and colleagues (2025) documented protocol asymmetries on Chatbot Arena, including 27 private Meta variants tested before Llama-4’s release, selective disclosure, and unequal access to arena data.
Cited in 1 location
-
arXiv:2510.22614
A large-scale null result from Koohestani and colleagues (2025) found that generic Platt scaling did not improve confidence-acceptance alignment across 24M interactions; a related study of 153 developers preferred color-coded indicators over numeric ones.
Cited in 1 location
-
arXiv:2604.07830
A Microsoft survey of 860 developers by Choudhuri (2026) appears in the same explorer synthesis and supports the direction of the practice, but contributes no cataloged figure about this mechanism.
Cited in 1 location
-
arXiv:2408.00440
A mining study by Laigner and colleagues (2024) examined 8000+ Stack Overflow questions and manually coded 628 of them.
Cited in 1 location
-
arXiv:2607.00533
A mixed-methods study of 448 professional developers by Choudhuri and colleagues (2026) found that task accountability was associated with lower odds of allowing AI to act on the developer’s behalf.
Cited in 1 location
-
arXiv:2603.11262
A null-result study by Williams and colleagues (2026) found that random selection outperformed all six evaluated state-of-the-art overfitting-detection tools in 71-96% of cases on curated realistic patch distributions.
Cited in 1 location
-
arXiv:2604.28138
A preprint by Wu and colleagues (2026) provides directional evidence that effect-aware checkpointing at these boundaries can avoid unnecessary snapshot work while preserving recovery semantics.
Cited in 1 location
-
arXiv:2111.11562
A process-calculus analysis by Tardieu and colleagues (2021) directionally establishes that these three properties jointly yield exactly-once effects when all observable application state, including entity state, queues, and orchestration progress, resides in the managed persistent store.
Cited in 2 locations
-
arXiv:2603.20625
A proof-of-concept experiment reported by Zheng (2026) demonstrates the semantic-rollback hazard and proposes a mitigation, but it does not measure the receipt-and-compensator design described here.
Cited in 1 location
-
arXiv:2502.15963
A qualitative study by Alami and colleagues (2025), based on 16 interviews and focus-group simulations, found that reciprocity and being seen by peers helped move code-quality responsibility from an individual concern toward collective accountability.
Cited in 1 location
-
arXiv:2607.09510
A second directional observational study by Zhao and colleagues (2026) examined 1,794 annotated command-line coding-agent trajectories containing 63,000+ steps.
Cited in 1 location
-
arXiv:2311.04850
A separate controlled intervention by Yang and colleagues (2023) showed that rephrasing and translation could bypass n-gram deduplication.
Cited in 1 location
-
arXiv:2604.13725
A separate empirical investigation by Feng and colleagues (2026) measured a more conditional result.
Cited in 1 location
-
arXiv:2602.11619
A single-author preprint by Mehta (2026) found a directional association between greater action-path divergence and poorer outcomes in ReAct-style agents on HotpotQA, with divergence becoming visible early enough to inform per-task review routing.
Cited in 1 location
-
arXiv:2102.00701
A single-organization case study by Schröder and colleagues (2021) examined search-based re-modularization at Adyen across 5.5M+ LOC.
Cited in 1 location
-
arXiv:2502.06215
A strong census by Zhou and colleagues (2025) examined 83 software-engineering benchmarks.
Cited in 1 location
-
arXiv:2001.08236
A survey by Chen and colleagues (2020) provides directional support for the selection mechanism: weighted-search solutions can be dominated by Pareto-search alternatives, while fixed weights can over-constrain search and create a burdensome weight-specification problem.
Cited in 1 location
-
arXiv:2404.00699
A survey by Ravaut and colleagues (2024) organized detection methods by access assumptions and contamination type, showing that their coverage differs.
Cited in 1 location
-
arXiv:2407.00125
A survey of 160 sources by Yu and colleagues (2024) found persistent discrepancies between injected fault menus and observed failures across six AI-system layers.
Cited in 1 location
-
arXiv:2112.11471
A survey synthesis by Lai and colleagues (2021) found that aggregate human-AI performance measures can hide opposing changes in over-reliance and under-reliance.
Cited in 1 location
-
arXiv:1506.08603
A systems technical report by Carbone and colleagues (2015) showed that asynchronous barrier snapshotting produced consistent snapshots without pausing the dataflow; for acyclic topologies, only operator state had to be persisted.
Cited in 1 location
-
arXiv:1709.02592
A theoretical result by Dürr and colleagues (2017) proves that a threshold policy is optimal for deterministic algorithms inside an adversarial single-machine model with unit test cost.
Cited in 1 location
-
arXiv:2603.16496
AdaMem (Yan et al. 2026), arXiv:2603.16496, carried in the same evidence record as the Cloudflare account. Its own retrieval route is semantic retrieval with conditional graph expansion rather than rank fusion, so it supports retaining a complementary semantic lane only.
Cited in 2 locations
-
arXiv:2507.11059
Adamenko et al. (2025), arXiv:2507.11059, SWE-MERA. [Directional evidence]
Cited in 1 location
-
arXiv:2108.13264
Agarwal and colleagues (2021) demonstrated this combination in multi-task reinforcement-learning suites with a handful of runs.
Cited in 1 location
-
arXiv:2406.13352
AgentDojo (Debenedetti et al. 2024, arXiv:2406.13352).
Cited in 1 location
-
arXiv:2510.04618
Agentic Context Engineering, arXiv:2510.04618.
Cited in 2 locations
-
arXiv:2605.11026
AgentShield deception-based detection (Rassul and Rashid 2026, arXiv:2605.11026).
Cited in 1 location
-
arXiv:2508.02736
AgentSight eBPF observability (Zheng et al. 2025, arXiv:2508.02736).
Cited in 1 location
-
arXiv:2502.11448
AGrail by Luo (2025) and Architecture Matters (2026) point in the same architectural direction without supplying a production effect size for this design.
Cited in 1 location
-
arXiv:2604.13018
AiScientist long-horizon engineering (Chen 2026), arXiv:2604.13018, carried by the companion record on durable artifact handoff.
Cited in 2 locations
-
arXiv:2410.06992
Aleithan et al. (2024), arXiv:2410.06992, SWE-Bench+. [Directional evidence]
Cited in 1 location
-
arXiv:2605.23023
AMBIPOM human-LLM co-planning (He 2026, arXiv:2605.23023).
Cited in 2 locations
-
arXiv:2406.10229
An empirical benchmark-variance study by Madaan and colleagues (2024) measured seed and checkpoint variability and found that choice log-likelihood detected weak signals better than discrete accuracy under the studied conditions.
Cited in 1 location
-
arXiv:2002.08035
An observational study by De-Arteaga and colleagues (2020) examined decisions at a child-maltreatment hotline.
Cited in 1 location
-
arXiv:1911.11915
Analytical work by Jayasekara and colleagues (2019) directionally supports deriving the interval from checkpoint cost and failure rate instead of accepting a stock default.
Cited in 1 location
-
arXiv:2607.07989
Another directional author-run evaluation by Xia and colleagues (2026) combined independent evaluators using confidence-aware aggregation and fed disagreement back into the judging process.
Cited in 1 location
-
arXiv:2605.02421
AOCI AI-oriented code indexing (Liu 2026), arXiv:2605.02421.
Cited in 1 location
-
arXiv:2409.09927
As a counterweight, a contested comparative evaluation by Samuel and colleagues (2024) found assumption violations, limited consistency across five methods, and failures involving modern post-trained models.
Cited in 1 location
-
arXiv:2510.04905
At the system level, a survey by Tao and colleagues (2025) identifies the boundary between retrieval-augmented code generation and long context as an open measurement problem.
Cited in 1 location
-
arXiv:2604.21131
Azarafrooz (2026) directionally reports that central session-bound gates miss cross-session and context-fragmented violations.
Cited in 1 location
-
arXiv:2605.23459
Badagi (2026), "AI Assurance," arXiv:2605.23459.
Cited in 1 location
-
arXiv:2505.20411
Badertdinov et al. (2025), arXiv:2505.20411, SWE-rebench, Nebius. [Directional evidence]
Cited in 1 location
-
DOI:10.1016/0005-1098(83
Bainbridge (198390046-8)) described this problem in industrial process control as an “irony of automation.” Its application to software-agent oversight is analogical rather than measured.
Cited in 1 location
-
DOI:10.1016/0005-1098(83)90046-8
Bainbridge, L. (1983). Ironies of Automation. Automatica 19(6), 775-779. doi:10.1016/0005-1098(83)90046-8. This is the single source behind the skill-retention paragraph. The evidence is pre-AI and industrial. Transfer to code review is analogical.
Cited in 1 location
-
arXiv:2206.15000
Barke, S., James, M.B., Polikarpova, N. (2022/2023). Grounded Copilot: How Programmers Interact with Code-Generating Models. PACMPL 7(OOPSLA1). arXiv:2206.15000. Grounded theory from 20 programmers supports the suggestion-level acceleration/exploration split. Transfer to autonomous-agent diffs is untested. This…
Cited in 1 location
-
arXiv:2511.04703
Bean, A.M. et al. (2025), Measuring what Matters, arXiv:2511.04703. [Directional evidence] Carried in the companion catalog as the construct-validity record; named inline for the construct-validity definition that opens this section.
Cited in 2 locations
-
arXiv:1907.07817
Bellm and colleagues (2019) explicitly proposed decision-reason logging, shared schedule-reporting schemas, and open-source schedulers in an Astro2020 position paper.
Cited in 1 location
-
arXiv:1905.02209
Bellm, E. C., et al. (2019), "The Zwicky Transient Facility: Surveys and Scheduler," PASP 131, 068003, arXiv:1905.02209.
Cited in 2 locations
-
arXiv:2605.00914
Bertalanič (2026), "Cost of Consensus," arXiv:2605.00914. Oracle gap up to 32.3 percentage points; plurality voting discards correct answers already present; grounding catches failures that agreement masks.
Cited in 2 locations
-
arXiv:2405.14782
Biderman and colleagues (2024) give directional support through a three-year operational account of the lm-evaluation-harness.
Cited in 1 location
-
arXiv:2602.07150
Bjarnason, Silva & Monperrus (2026). On Randomness in Agentic Evals. arXiv:2602.07150. The 60,000-trajectory result: 2.2 to 6.0 pp single-run spread, SD above 1.5 pp at temperature 0, early-token divergence.
Cited in 1 location
-
arXiv:2604.14723
Bounded Autonomy for Enterprise AI (Sohail 2026, arXiv:2604.14723).
Cited in 2 locations
-
arXiv:2103.03098
Bouthillier et al. (2021). Accounting for Variance in Machine Learning Benchmarks. MLSys 2021. arXiv:2103.03098. The ~51x result and the fixed-seed critique.
Cited in 1 location
-
arXiv:2108.03758
Braun and colleagues (2021) report directional evidence from two action-research cycles in one industrial context.
Cited in 1 location
-
arXiv:2407.21787
Brown and colleagues (2024) found that this relationship held over four orders of magnitude; on SWE-bench Lite, coverage rose from 15.9% at one sample to 56% at 250 samples.
Cited in 1 location
-
arXiv:2102.09692
Buçinca, Z., Malaya, M.B., Gajos, K.Z. (2021). To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. PACM HCI 5(CSCW1), 188. arXiv:2102.09692. Cognitive forcing functions reduced overreliance where explanations failed (N=199).
Cited in 1 location
-
arXiv:2501.10711
Cao and colleagues (2025) conducted an observational profile of 274 code-related benchmarks and reported widespread gaps in basic quality assurance.
Cited in 1 location
-
arXiv:2010.06595
Card et al. (2020). With Little Power Comes Great Responsibility. EMNLP 2020. arXiv:2010.06595. GLUE underpowering and effect-size exaggeration.
Cited in 1 location
-
arXiv:2112.01298
Cavalcante Siebert, L., et al. (2023). Meaningful human control: actionable properties for AI system development. AI and Ethics 3, 241-255. arXiv:2112.01298. Supports responsibility commensurate with ability and authority to control, the framework's third actionable property.
Cited in 1 location
-
arXiv:2003.00365
Cayci, S., Eryilmaz, A., & Srikant, R. (2020), "Budget-Constrained Bandits over General Cost and Reward Distributions," arXiv:2003.00365. Asymptotic guarantee under stated moment conditions.
Cited in 2 locations
-
arXiv:2503.13657
Cemri, M., et al. 2025. "Why Do Multi-Agent LLM Systems Fail?" arXiv:2503.13657. MAST failure taxonomy and benchmark-framework interventions.
Cited in 2 locations
-
arXiv:2401.13138
Chan, A., et al. (2024), "Visibility into AI Agents," ACM FAccT 2024, arXiv:2401.13138.
Cited in 1 location
-
arXiv:2502.15292
Chang, J., et al. (2025), "BugCerberus: Bridging Bug Localization and Issue Fixing," arXiv:2502.15292. (Per-hierarchy-level specialization.)
Cited in 1 location
-
arXiv:2511.12884
Chatlatanagulchai, W., et al. (2025), "Agent READMEs: An Empirical Study of Context Files for Agentic Coding," arXiv:2511.12884.
Cited in 1 location
-
arXiv:2309.13007
Chen et al. (2023), "ReConcile," arXiv:2309.13007.
Cited in 1 location
-
arXiv:2502.17521
Chen, S. et al. (2025), arXiv:2502.17521, static-to-dynamic benchmark survey. [Directional evidence]
Cited in 1 location
-
arXiv:2507.08149
Chen, V., Talwalkar, A., Brennan, R., Neubig, G. (2025). Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows. arXiv:2507.08149. In the controlled study, users' failure to understand agent behavior, not capability, limited broader adoption.
Cited in 1 location
-
arXiv:2503.09089
Chen, Z., et al. (2025), "LocAgent: Graph-Guided LLM Agents for Code Localization," arXiv:2503.09089.
Cited in 1 location
-
arXiv:2005.13109
Choudhury and colleagues (2020) evaluated SCoBA on multi-drone delivery and conveyor sorting, computing worker policies separately before coordinating conflicts.
Cited in 1 location
-
arXiv:2503.12029
Chun, Jina, et al. 2025. "Is Multi-Agent Debate the Silver Bullet?" arXiv:2503.12029. Audit-retained strong evidence on debate performance against task baselines; reports inference cost across debate variants and leaves the single-model cost comparison open.
Cited in 2 locations
-
arXiv:1806.08295
Colas, Sigaud & Oudeyer (2018). How Many Random Seeds? arXiv:1806.08295. The pilot floor, the bootstrap false-positive rate at small N, the DDPG case, and the preference for Welch's t-test.
Cited in 1 location
-
arXiv:2004.12832
ColBERT (Khattab and Zaharia 2020), arXiv:2004.12832.
Cited in 2 locations
-
arXiv:2406.19228
Controlled calculator and embodied-planning experiments by Sun and colleagues (2024) provide directional evidence that detection followed by reflection can improve recovery from faulty tool outputs.
Cited in 1 location
-
arXiv:2307.03172
Controlled multi-document question-answering and key-value retrieval experiments by Liu and colleagues (2023) measured a U-shaped sensitivity curve: models used relevant material less reliably when it appeared near the middle of a long context.
Cited in 1 location
-
arXiv:2601.07978
Cost-and-Accuracy study (Wolff & Bennati 2026), arXiv:2601.07978.
Cited in 2 locations
-
arXiv:2512.07094
Cruz (2025) describes one state-machine-gated supervision case in which only legal transitions could execute.
Cited in 1 location
-
arXiv:1810.01963
Decima, Mao, H., et al. (2018), "Learning Scheduling Algorithms for Data Processing Clusters," arXiv:1810.01963. Measures learned scheduling under stochastic arrivals, not fixed-arrival replay; the replay transfer is the synthesis's.
Cited in 1 location
-
arXiv:2505.08638
Deshpande, D., et al. (2025), "TRAIL: Trace Reasoning and Agentic Issue Localization," Patronus AI, arXiv:2505.08638. (Best evaluated model localized issues in 11% of 148 expert-annotated traces; a 2025 capability snapshot.)
Cited in 1 location
-
arXiv:2212.08385
Di Pompeo and Tucci (2022) conducted a controlled comparison of 3 algorithms x 3 budgets x 31 repetitions, with hypothesis testing, on model-based refactoring.
Cited in 1 location
-
arXiv:2509.25257
Directional benchmark evidence from Shah and colleagues (2025) supports this division of query traffic.
Cited in 1 location
-
arXiv:2602.05665
Directional evidence (boundary citation): Graph-based Agent Memory survey (Yang 2026), arXiv:2602.05665. Cited as the counterweight, not support.
Cited in 1 location
-
arXiv:2607.28802
Directional evidence for interaction-centered failure localization, with strong evidence for label reproducibility in the evaluated set: Raj, H., et al. (2026), "Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures," arXiv:2607.28802. The strongest automated judge reached Cohen's…
Cited in 1 location
-
arXiv:2604.12262
Directional evidence for this developed practice: CascadeDebate (Chang 2026), arXiv:2604.12262. Its confidence-gated cascade comparison strongly supports the narrower companion-catalog claim in the tested setting; transfer to production software-agent routing remains directional.
Cited in 2 locations
-
arXiv:2606.01326
Directional evidence from Hrubec and Cito (2026) found that token savings from minification accompanied lower SWE-bench Verified resolution.
Cited in 1 location
-
arXiv:2407.05337
Directional evidence from Mulone and colleagues (2024) shows per-task provenance combined with rollback and re-execution restoring the affected subgraph in the StreamFlow scientific workflow management system.
Cited in 1 location
-
arXiv:2605.23311
Directional evidence from Yang, Ke and colleagues (2026) demonstrates the distinction in a LangGraph reconstruction: commitment-blocking rejected an alignment-legal restore that would have orphaned two committed consumers.
Cited in 1 location
-
arXiv:2207.11016
Directional evidence: Formica, Fan & Menghi (2022). Search-based Software Testing Driven by Automatically Generated and Manually Defined Fitness Functions. arXiv:2207.11016.
Cited in 1 location
-
arXiv:2403.03185
Directional evidence: Laidlaw, C., et al. (2024). Correlated Proxies. arXiv:2403.03185.
Cited in 1 location
-
arXiv:2406.16008
Directional experiments by Hsieh and colleagues (2024) attributed lost-in-the-middle behavior to intrinsic positional bias and found that calibrating attention toward relevance improved long-context retrieval-augmented generation.
Cited in 1 location
-
arXiv:2407.06738
Directional formal evidence from Veresov and colleagues (2024) defines failure transparency through observational explainability and proves the property for a small-step model of Flink’s protocol when observers see only committed state monotonically.
Cited in 1 location
-
arXiv:2602.05891
Directional practitioner evidence from Zheng and colleagues (2026) found that submission order, contest selection, and repeated runs changed Codeforces-based Elo rankings.
Cited in 1 location
-
arXiv:2103.00033
Directional systems evidence from Burckhardt and colleagues (2021) supports deterministic control flow over logged history as the basis of replay recovery.
Cited in 2 locations
-
arXiv:2502.14847
Directional work on communication attacks by He (2025), G-Safeguard by Wang (2025), and the MPAC protocol (2026) points toward treating communication paths and permission boundaries as security controls.
Cited in 1 location
-
arXiv:2502.11127
Directional work on communication attacks by He (2025), G-Safeguard by Wang (2025), and the MPAC protocol (2026) points toward treating communication paths and permission boundaries as security controls.
Cited in 1 location
-
arXiv:2604.09744
Directional work on communication attacks by He (2025), G-Safeguard by Wang (2025), and the MPAC protocol (2026) points toward treating communication paths and permission boundaries as security controls.
Cited in 1 location
-
arXiv:1909.03004
Dodge and colleagues (2019) showed empirically that best-of-B is a biased order statistic and that published method rankings can reverse when tuning budgets are equalized.
Cited in 1 location
-
arXiv:2002.06305
Dodge et al. (2020). Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping. arXiv:2002.06305. The 2,100-trial seed result, scoped to small-data fine-tuning of pretrained encoders.
Cited in 1 location
-
arXiv:1809.01448
Dror & Reichart (2018). Appendix, Recommended Statistical Significance Tests for NLP Tasks. arXiv:1809.01448. The per-metric test selection; directional, so the table is guidance, not a measured result.
Cited in 1 location
-
arXiv:2604.21229
EngramaBench (Acuna 2026), arXiv:2604.21229.
Cited in 1 location
-
arXiv:2407.02883
evaluation-design material from the code-retrieval source corpus. CoIR (Li et al. 2024, arXiv:2407.02883) supplies code-retrieval tasks and metrics, but it does not test the no-tool baseline, tool-access control, or ground-truth-tautology checks recommended here.
Cited in 1 location
-
arXiv:2305.07722
Fok, R., Weld, D.S. (2023/2024). In Search of Verifiability: Explanations Rarely Enable Complementary Performance in AI-Advised Decision Making. AI Magazine 45(3). arXiv:2305.07722. Across the XAI-reliance literature, explanations help only insofar as they enable verification.
Cited in 1 location
-
arXiv:2412.01007
For code localization, Suresh and colleagues (2024) found that consistency-filtered contrastive data and curriculum hard-negative mining yielded state-of-the-art code retrieval and reranking on real GitHub issues.
Cited in 1 location
-
DOI:10.1177/001316446002000104
Foundational method: Cohen, J. (1960). A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20(1), 37-46. DOI: 10.1177/001316446002000104. Primary source for Cohen's kappa, the agreement statistic the chapter defines; a standard method rather than a catalog evidence item.
Cited in 1 location
-
DOI:10.1037/h0031619
Foundational method: Fleiss, J. L. (1971). Measuring Nominal Scale Agreement Among Many Raters. Psychological Bulletin 76(5), 378-382. DOI: 10.1037/h0031619. Primary source for Fleiss's kappa, the agreement statistic the chapter defines; a standard method rather than a catalog evidence item.
Cited in 1 location
-
DOI:10.1007/BF02295996
Foundational method: McNemar, Q. (1947). Psychometrika 12(2), 153-157. DOI: 10.1007/BF02295996. The paired pass/fail test used in the binary-outcome row.
Cited in 1 location
-
DOI:10.1086/266577
Foundational method: Scott, W. A. (1955). Reliability of Content Analysis: The Case of Nominal Scale Coding. Public Opinion Quarterly 19(3), 321-325. DOI: 10.1086/266577. Primary source for Scott's pi, the agreement statistic the chapter defines; a standard method rather than a catalog evidence item.
Cited in 1 location
-
arXiv:2008.00842
Fragkoulis, Carbone, Kalavri & Katsifodimos (2020). A Survey on the Evolution of Stream Processing Systems. The VLDB Journal (2024). arXiv:2008.00842. (Exactly-once recovery arrived only when state became a first-class managed runtime artifact; implicit-state systems had lossy recovery.)
Cited in 1 location
-
arXiv:2408.15204
Gligorić and colleagues (2024) found that statistically valid confidence-driven allocation outperformed both all-LLM annotation and uniform-random human allocation.
Cited in 1 location
-
arXiv:2109.05067
Green, B. (2022). The Flaws of Policies Requiring Human Oversight of Government Algorithms. Computer Law & Security Review 45, 105681. arXiv:2109.05067. Compares 41 policies with the human-computer interaction record and shifts the burden of proof to the deploying institution. The negative claim transfers only where…
Cited in 1 location
-
arXiv:2009.08366
Guo and colleagues (2020) likewise provide strong evidence that structural and data-flow signals support code retrieval, while Ye and colleagues (2020) support the direction of this practice without supplying a cataloged magnitude.
Cited in 1 location
-
arXiv:2006.05265
Guo and colleagues (2020) likewise provide strong evidence that structural and data-flow signals support code retrieval, while Ye and colleagues (2020) support the direction of this practice without supplying a cataloged magnitude.
Cited in 1 location
-
arXiv:2410.09247
Haimes, Wenner et al. (2024), arXiv:2410.09247, Apart Research. [Strong evidence]
Cited in 1 location
-
arXiv:2604.25849
Hanlin (Zhou) et al. (2026). ADEMA: A Knowledge-State Orchestration Architecture for Long-Horizon Knowledge Synthesis with LLM Agents. arXiv:2604.25849. (In a fixed 60-run matrix, removing checkpoint/resume produced the only invalid run.)
Cited in 1 location
-
arXiv:2505.24201
He and colleagues (2025) reported graph-level anomaly detection with explainable root-cause attribution in two case studies involving Magentic-One and an email assistant.
Cited in 1 location
-
arXiv:2508.13144
Heineman and colleagues (2025) analyzed 30 benchmarks and 375 models.
Cited in 1 location
-
arXiv:1709.06560
Henderson et al. (2017). Deep Reinforcement Learning that Matters. AAAI 2018. arXiv:1709.06560. Five-seed splits and unreported researcher degrees of freedom.
Cited in 1 location
-
arXiv:2404.06654
Hsieh, C.-P., et al. (2024), "RULER: What's the Real Context Size of Your Long-Context Language Models?" COLM 2024, arXiv:2404.06654.
Cited in 1 location
-
arXiv:2310.01798
Huang, J., et al. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. arXiv:2310.01798.
Cited in 1 location
-
arXiv:2605.13850
Huang, Jia, and Joey Tianyi Zhou. 2026. "A Two-Dimensional Framework for AI Agent Design Patterns: Cognitive Function and Execution Topology." arXiv:2605.13850. Directional taxonomy of topology choices.
Cited in 2 locations
-
arXiv:2406.14497
In a benchmark spanning five retrieval sources and three task classes, Wang and colleagues (2024) found bottlenecks at both ends: retrievers struggled when lexical overlap was limited, while generators sometimes failed to use supplied context under constrained windows or weak context-integration ability.
Cited in 1 location
-
arXiv:2107.03374
In a code-generation evaluation study, Chen and colleagues (2021) introduced this combinatorial estimator and showed that it is exactly unbiased for the probability that at least one of k draws succeeds, while the plug-in alternative is biased high.
Cited in 1 location
-
arXiv:2502.13295
In a comparative chess study, Bondarenko and colleagues (2025) observed that reasoning models, including o1-preview and DeepSeek-R1, hacked an infeasible task without being nudged toward that strategy, while non-reasoning models required explicit prompting that normal play would not work.
Cited in 1 location
-
arXiv:2407.10457
In a comparative evaluation across model sizes and tasks, Song and colleagues (2024) found a consistent greedy-versus-sampling gap large enough to reorder model comparisons.
Cited in 1 location
-
arXiv:2406.10162
In a constructed curriculum study, Denison and colleagues (2024) found that models rewarded for easy forms of gaming generalized to rarer and more consequential behavior, including zero-shot reward-function rewriting.
Cited in 1 location
-
arXiv:2302.07248
In a controlled code-completion study, Vasconcelos and colleagues (2024) found that edit-likelihood highlighting reduced task time and directed edits toward highlighted tokens.
Cited in 1 location
-
arXiv:2310.18497
In a controlled comparison of 360 scheduling instances, Handley and colleagues (2024) found that the heuristic failed on seven instances.
Cited in 1 location
-
arXiv:2510.16786
In a controlled SWE-bench study with three state-of-the-art models, Gao and Peng (2025) found that 75th-percentile caps reduced costs 24-68% with minimal solve-rate impact; dynamic extension reduced costs a further 12-24% relative to fixed caps.
Cited in 1 location
-
arXiv:2605.02273
In a descriptive observational study of 33,596 agent-authored GitHub pull requests, Duma and colleagues (2026) found that 61.4% had no recorded review.
Cited in 1 location
-
arXiv:2510.04886
In a directional author-run evaluation, Banerjee and colleagues (2025) found that hierarchical context leveling with consensus voting outperformed all-at-once, step-by-step, and binary-search attribution.
Cited in 1 location
-
arXiv:2601.20103
In a directional benchmark study of 517 human-verified trajectories spanning a 54-category exploit taxonomy, Deshpande and colleagues (2026) found that contrastive evaluation improved detection relative to isolated classification.
Cited in 1 location
-
arXiv:2507.03870
In a directional black-box evaluation, Anand and colleagues (2025) used an independent planner to separate systemic agent errors from environment errors and reported detecting more total and unique errors than the evaluated prior methods across their domains.
Cited in 1 location
-
arXiv:2602.02475
In a directional evaluation of 115 annotated failed trajectories, Barke and colleagues (2026) found that constraint-violation logs improved step localization and attribution over judging raw traces across structured application workflows, incident management, and open-ended web and file tasks.
Cited in 1 location
-
arXiv:2403.04132
In a large human-preference evaluation study, Chiang and colleagues (2024) used Bradley-Terry aggregation with published confidence intervals, active pair sampling, and anomalous-voter screening.
Cited in 1 location
-
arXiv:1709.09500
In a multi-dataset NLP study, Dror and colleagues (2017) demonstrated that partial-conjunction testing controls family-wise error and supports statistically valid claims about the minimum number of datasets showing an advantage.
Cited in 1 location
-
arXiv:2502.03461
In a platinum-grading study of fifteen popular benchmarks, Vendrow and colleagues (2025) found that cleaning label errors and ambiguity exposed persistent frontier-model failures on simple tasks that the dirty sets had obscured.
Cited in 1 location
-
arXiv:2412.15584
In a randomized experiment with 400 lay participants, Bo and colleagues (2024) reported that the disclaimer shifted both over-reliance and under-reliance in the desired direction across single-shot advice tasks involving LSAT logic and image-based estimation.
Cited in 1 location
-
arXiv:2303.12570
In a repository-level code-completion comparison, Zhang and colleagues (2023) found that iterative retrieval and generation beat single-shot RAG and improved over in-file-only baselines by >10% on RepoEval.
Cited in 1 location
-
arXiv:2111.11628
In a separate historical validation, Claudet and colleagues (2021) encoded a heuristic for balancing mission satisfaction in Delta-MILP and evaluated schedules on the most oversubscribed weeks of 2016 and 2018 Deep Space Network demand.
Cited in 2 locations
-
arXiv:2511.16842
In a separate study across nine benchmarks, Truong and colleagues (2025) combined item-statistic outlier detection with model-assisted triage and achieved up to 84% precision when identifying problematic questions.
Cited in 1 location
-
arXiv:2402.06627
In a simulated deployment study, Pan and colleagues (2024) found that feedback loops drove in-context reward hacking at test time.
Cited in 1 location
-
arXiv:2404.16966
In an empirical analysis of major benchmarks, Ailem and colleagues (2024) found non-random performance correlations across prompts; accounting for those correlations enlarged naive standard errors and changed model rankings.
Cited in 1 location
-
arXiv:2510.01367
In an evaluation study, Wang and colleagues (2025) showed that truncated-reward AUC detected implicit reward hacking without interpreting chain-of-thought content.
Cited in 1 location
-
arXiv:2403.10059
In code-completion experiments, Wu and colleagues (2024) found that threshold-based selective retrieval beat always-retrieving and greedy selection on RepoEval and CrossCodeEval.
Cited in 1 location
-
arXiv:2401.12187
In controlled summarization experiments, Ramé and colleagues (2024) found that weight-averaged reward models improved reliability under distribution shift and resistance to preference noise in best-of-N selection and reinforcement learning.
Cited in 1 location
-
arXiv:2403.07008
In GPT-4 experiments, Boyeau and colleagues (2024) obtained unbiased estimates with valid confidence intervals and effectively increased human sample size by up to 50%.
Cited in 1 location
-
arXiv:2604.23459
In tested configurations, Hagag (2026) found up to 3.8x attack-success variance across topologies for the same task, with untrusted input propagating across agent boundaries.
Cited in 2 locations
-
arXiv:2302.12173
Indirect prompt injection (Greshake et al. 2023, arXiv:2302.12173); companion material named in the same synthesis: OWASP LLM Top 10.
Cited in 1 location
-
arXiv:2305.10160
Jacovi and colleagues (2023) provide directional support in a policy paper arguing for publisher-side prevention through encryption, licensing, exclusion guarantees, and regenerated instances.
Cited in 1 location
-
arXiv:2403.07974
Jain et al. (2024), arXiv:2403.07974, LiveCodeBench, ICLR 2025. [Strong evidence]
Cited in 1 location
-
arXiv:2604.01527
Jha, Paltenghi, Maddila, Murali, Ugare, and Chandra (2026), REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage, arXiv:2604.01527; the resulting benchmark is Harvest. [Directional evidence] The title and benchmark name were verified against the arXiv record on 2026-07-29; earlier…
Cited in 1 location
-
arXiv:2602.19843
Jia, J., et al. 2026. "MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems." arXiv:2602.19843. Synthetic fault injection across three architectures.
Cited in 2 locations
-
arXiv:2503.22424
Jiang and colleagues (2025) provide directional evidence from CoSIL that query-time module and function graph expansion can support issue localization without a pre-built index.
Cited in 1 location
-
arXiv:2604.13108
Jin (2026) found directional evidence that architecture descriptors reduced navigation work, while detecting no significant comprehension difference among S-expression, JSON, YAML, and Markdown.
Cited in 1 location
-
DOI:10.1145/3772370
Johnson, B., et al. (2026). "Facilitating Trust in AI-assisted Software Tools." ACM Transactions on Software Engineering and Methodology. DOI:10.1145/3772370. Eighteen interviews and a 368-response survey identify factors shaping trust in software tools; autonomous patch-gate transfer remains untested.
Cited in 1 location
-
arXiv:2508.19461
Kale, N., et al. (2025). Reliable Weak-to-Strong Monitoring of LLM Agents. arXiv:2508.19461. Targeted escalation adds approximately 15% TPR at FPR=0.01. Hybrid hierarchical-sequential scaffolds let weaker models monitor stronger agents. The paper reports that scaffolding mattered more than monitor awareness; the…
Cited in 1 location
-
arXiv:2604.23366
Kamelhar (2026), GSAR grounded consensus, arXiv:2604.23366.
Cited in 1 location
-
arXiv:2407.01502
Kapoor, Stroebl, Siegel, Nadgir, and Narayanan (2024), AI Agents That Matter, arXiv:2407.01502, a re-evaluation study.
Cited in 3 locations
-
arXiv:2606.17182
Khan (2026) used TLA+ specifications and TLC counter-examples to exhibit all four anomaly classes and showed directionally that isolation mechanisms exclude them under deterministic-generation semantics.
Cited in 1 location
-
arXiv:2602.09937
Kim, T., et al. 2026. "Why Do AI Agents Systematically Fail at Cloud Root Cause Analysis?" arXiv:2602.09937. OpenRCA evidence on failure persistence across capability tiers and protocol intervention.
Cited in 1 location
-
arXiv:2503.05860
Koohestani and colleagues (2025) provide directional evidence from a review-and-tooling contribution that catalogued 204 AI4SE benchmarks with overlapping constructs.
Cited in 1 location
-
arXiv:2604.13120
Kumar, et al. 2026. AgentForge. arXiv:2604.13120. Directional evidence on multi-agent design.
Cited in 2 locations
-
arXiv:2505.06120
Laban, P., et al. (2025), "LLMs Get Lost In Multi-Turn Conversation," arXiv:2505.06120.
Cited in 1 location
-
arXiv:2103.00170
Laigner, Zhou, Vaz Salles et al. (2021). Data Management in Microservices: State of the Practice, Challenges, and Research Directions. PVLDB 14(13). arXiv:2103.00170. (SLR + repo analysis + 120+ practitioner survey: hand-rolled sagas and convention-managed consistency are where reliability failures concentrate.)
Cited in 1 location
-
arXiv:1503.07170
Lampoudi, S., Saunders, E., & Eastman, J. (2015), "An Integer Linear Programming Solution to the Telescope Network Scheduling Problem," arXiv:1503.07170. Supports validating a scheduling kernel on instances with known optima.
Cited in 3 locations
-
arXiv:2407.12220
Leech and colleagues (2024) provide a directional taxonomy of 44 questionable machine-learning research practices and examples of how individually defensible decisions could accumulate in one favorable direction.
Cited in 1 location
-
arXiv:2411.03538
Leng, Q., et al. (2024), "Long Context RAG Performance of Large Language Models," Databricks Mosaic Research, arXiv:2411.03538.
Cited in 1 location
-
arXiv:2502.01534
Li (2025), "Preference Leakage," arXiv:2502.01534.
Cited in 1 location
-
arXiv:2402.07632
Li and colleagues (2024) found that overconfident AI produced misuse and unconfident AI produced disuse.
Cited in 1 location
-
arXiv:2606.00655
Li, et al. 2026. SIMAS. arXiv:2606.00655. Audit-retained strong evidence on token cost, non-monotonic scaling, and debate versus self-correction.
Cited in 1 location
-
arXiv:2503.04359
Li, J., et al. (2025), "LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding," arXiv:2503.04359.
Cited in 1 location
-
arXiv:2502.02743
Li, Y. (2025), "LLM Bandit: Cost-Efficient LLM Generation via Preference-Conditioned Dynamic Routing," arXiv:2502.02743. Single-author preprint, benchmark evidence.
Cited in 2 locations
-
arXiv:2606.26979
Lin and colleagues (2026) provide directional evidence that static-structure annotations improve run-to-run stability more than raw capability on a strong Codex baseline.
Cited in 1 location
-
arXiv:2406.07003
Liu and colleagues (2024) compared dependence-graph retrieval with sequence-similarity RAG baselines for repository-level completion.
Cited in 1 location
-
arXiv:2306.03091
Liu, T., et al. (2023), "RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems," ICLR 2024, arXiv:2306.03091. (R/C/P decomposition localizes failure to retrieval vs completion vs pipeline stage.)
Cited in 1 location
-
arXiv:2408.03910
Liu, X., et al. (2024), "CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases," ICLR 2025, arXiv:2408.03910.
Cited in 1 location
-
arXiv:2305.01210
Liu, Xia, Wang & Zhang (2023), arXiv:2305.01210, HumanEval+ / EvalPlus, NeurIPS 2023. [Strong evidence]
Cited in 1 location
-
arXiv:2508.13143
Lu, R., Li, Y., Huo, Y. (2025), "Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks," arXiv:2508.13143. (Phase-aligned classification; ~50% completion base rate on 34 programmable tasks with off-the-shelf frameworks; not an estimate for production repositories.)
Cited in 1 location
-
arXiv:2502.13965
Luo, et al. 2025. Autellix. arXiv:2502.13965. Directional scheduling and orchestration example.
Cited in 1 location
-
arXiv:2509.08682
Ma, G., et al. (2025), "Automatic Failure Attribution and Critical Step Prediction Method for Multi-Agent Systems Based on Causal Inference," arXiv:2509.08682 (CDC-MAS). (36.2% step-level against under 15% for prior methods; counterfactually validated fixes worth an average 22.4% task success.)
Cited in 1 location
-
arXiv:2406.01422
Ma, Y., et al. (2024), "Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration," arXiv:2406.01422.
Cited in 1 location
-
arXiv:2403.17927
MAGIS (Tao 2024), arXiv:2403.17927.
Cited in 1 location
-
arXiv:2507.03525
Manheim and Homewood (2025) propose this distinction in a framework and maturity model that maps policy language onto technically feasible supervision mechanisms.
Cited in 1 location
-
arXiv:2510.02557
Masters, et al. 2025. Manager Agent research challenge. arXiv:2510.02557. Directional planning and scheduling example.
Cited in 3 locations
-
arXiv:2603.25764
Mehta, A. (2026). Confident and Wrong: Silent Semantic Failures in Coding Agents. arXiv:2603.25764. Analysis of 1,750 trajectories across 50 tasks and four models; single-author, limited-sample observational finding. Supports the companion-only entry on scoring verified resolution separately from submission.
Cited in 2 locations
-
arXiv:2606.22936
Mehta’s follow-up (2026) found that its commitment signal could not separate committed-correct from committed-wrong questions.
Cited in 1 location
-
arXiv:2605.03354
Memory Circuit Analysis (Mao et al. 2026), arXiv:2605.03354. (Stage-level diagnostic localizing a silent failure to the responsible memory operation; extraction/retention/retrieval fail silently behind fluent answers.)
Cited in 1 location
-
arXiv:2606.01138
memorywire (Munirathinam 2026), arXiv:2606.01138, carried by the companion record on memory portability.
Cited in 2 locations
-
arXiv:2411.00640
Miller (2024). Adding Error Bars to Evals. arXiv:2411.00640 (Anthropic). The 13.2% to 7.5% minimum-detectable-effect example and the response-count lever.
Cited in 2 locations
-
arXiv:2502.02649
Mitchell and colleagues (2025) contest the development of fully autonomous agents and argue that risk grows as control is ceded.
Cited in 1 location
-
arXiv:2401.00595
Mizrahi et al. (2024), TACL, arXiv:2401.00595. Individual templates reversed some model comparisons even where aggregate comparisons were more stable.
Cited in 2 locations
-
arXiv:2306.04930
Mozannar and colleagues (2023) retrospectively evaluated selective display using interaction data from 535 programmers using GitHub Copilot.
Cited in 1 location
-
arXiv:2503.01935
MultiAgentBench (Zhu 2025, arXiv:2503.01935).
Cited in 1 location
-
arXiv:2204.07210
Nadeem & Malik (2022). A Case for Microservices Orchestration Using Workflow Engines. ICSE-NIER. arXiv:2204.07210.
Cited in 1 location
-
arXiv:1810.04815
Naghib, E., et al. (2019), "A Framework for Telescope Schedulers: With Applications to the Large Synoptic Survey Telescope," AJ 157, 151, arXiv:1810.04815. Framework and simulation evidence.
Cited in 2 locations
-
arXiv:2607.09682
Nakrani (2026) reported a validation of this construction using randomized mutation trials.
Cited in 1 location
-
DOI:10.1109/TSE.2022.3174092
Nine methodologically material works were admitted after title-and-abstract screening, including the SEGRESS reporting guideline ([Kitchenham et al.
Cited in 1 location
-
arXiv:2602.11988
Null or conflicting result: Gloaguen, T., Mündler, N., Müller, M., Raychev, V., Vechev, M. (2026), "Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?" arXiv:2602.11988.
Cited in 1 location
-
arXiv:2607.27250
Null or conflicting result: Khatri, P. (2026), "Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories," arXiv:2607.27250. Across 288 runs, equivalence tests bounded correctness effects to no more than roughly 10 to 15 percentage points in the studied conditions.
Cited in 1 location
-
arXiv:2208.09827
Null or conflicting result: Zhang, Soto & Markl (2022). A Survey on Transactional Stream Processing. The VLDB Journal. arXiv:2208.09827. (Carried as the negative result behind the engine choice; it is about deterministic data pipelines and says nothing about agents.)
Cited in 1 location
-
arXiv:2604.22446
One explorer synthesis points in this direction across “From Skills to Talent” (2026), a 2-D agent design-pattern framework (2026), and Wu’s AutoGen (2023).
Cited in 1 location
-
arXiv:2308.08155
One explorer synthesis points in this direction across “From Skills to Talent” (2026), a 2-D agent design-pattern framework (2026), and Wu’s AutoGen (2023).
Cited in 1 location
-
arXiv:2607.08740
One position paper by Quinto and colleagues (2026) extends the model to LLM workflows through the derive and infer distinction; it establishes a proposed design, not an empirically validated recovery rate.
Cited in 1 location
-
arXiv:2608.01507
Oskooei, A. R., et al. 2026. "Deep Agentic Search for Repository-Level Code Question Answering: An Empirical Study." arXiv:2608.01507. Strong evidence for the SWE-QA comparison and its measured handoff failures; broader topology transfer is directional.
Cited in 1 location
-
arXiv:2410.14684
Ouyang, S., et al. (2024), "RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph," arXiv:2410.14684.
Cited in 1 location
-
arXiv:2308.02828
Ouyang, Zhang, Harman & Wang (2023). An Empirical Study of the Non-determinism of ChatGPT in Code Generation. arXiv:2308.02828. Nominally identical requests, different programs and outcomes.
Cited in 1 location
-
arXiv:2104.01146
Overeem and colleagues (2021) conducted a grounded-theory study involving 25 engineers and 19 production event-sourced systems.
Cited in 1 location
-
arXiv:2605.24060
Panthi & Abdelfattah (2026), "Same Ranking, Different Winner," arXiv:2605.24060. The 115-case stratified validation protocol, agreement measurements, overlap adjudication, and scaled contested-case evaluation.
Cited in 4 locations
-
arXiv:2203.00013
Parazin, B., et al. (2022), "Foraging with MUSHROOMS: A Mixed-integer Linear Programming Scheduler for Multimessenger Target of Opportunity Searches with the Zwicky Transient Facility," ApJ 935, 87, arXiv:2203.00013.
Cited in 1 location
-
arXiv:2304.01315
Patterson and colleagues (2023) provide directional methodological evidence, supported by demonstrations in reinforcement-learning research, for a pilot-then-precommit-then-confirm discipline.
Cited in 1 location
-
arXiv:2605.03409
Perera (2026) and Chang (2025) provide directional support for compensation-oriented recovery without supplying a result that establishes its effectiveness.
Cited in 1 location
-
arXiv:2503.11951
Perera (2026) and Chang (2025) provide directional support for compensation-oriented recovery without supplying a result that establishes its effectiveness.
Cited in 1 location
-
arXiv:2308.11696
Perlitz and colleagues (2023) introduced DIoR to quantify the decision-flip probability associated with reductions.
Cited in 1 location
-
arXiv:2211.03622
Perry, Srivastava, Kumar, Boneh (2022/2023), "Do Users Write More Insecure Code with AI Assistants?", ACM CCS 2023, arXiv:2211.03622.
Cited in 2 locations
-
arXiv:2602.01146
PersistBench from Pulipaka (2026) supports only the need for a forgetting lifecycle; it is a memory-safety benchmark rather than a study of consolidation schedules.
Cited in 1 location
-
arXiv:2506.17001
PersonalAI KG comparison (Menschikov), arXiv:2506.17001, with a contested practitioner thread named in the same synthesis.
Cited in 1 location
-
arXiv:2402.14992
Polo and colleagues (2024) found that 100 IRT-curated items reconstructed full MMLU scores within ~2 points at ~1% of the cost.
Cited in 1 location
-
arXiv:2512.10218
Prathifkumar, Saji Mathews & Nagappan (2025), arXiv:2512.10218, Waterloo. [Directional evidence]
Cited in 1 location
-
arXiv:2605.08112
Product context and coding-agent decision compliance (Dillon 2026), arXiv:2605.08112, carried by the companion record on the tribal-knowledge substrate. Its compliance figure is not used in the prose.
Cited in 2 locations
-
arXiv:2507.19942
Prometheus (Pan, H., et al. 2025), arXiv:2507.19942, carried by the companion record on persisting explored context. No figure quoted here.
Cited in 2 locations
-
arXiv:2312.06893
Psarakis et al. (2023). Styx: Transactional Stateful Functions on Streaming Dataflows. SIGMOD line. arXiv:2312.06893. (Execution caching keyed by invocation identity gives exactly-once composition across call graphs.)
Cited in 1 location
-
arXiv:2406.07155
Qian, et al. 2024. MacNet. arXiv:2406.07155. Directional evidence on network scaling.
Cited in 1 location
-
DOI:10.1145/3607179
Rahman, M. M., et al. (2023), "A Systematic Review of Automated Query Reformulations in Source Code Search," ACM Transactions on Software Engineering and Methodology, DOI:10.1145/3607179. The review supports treating query formation as a distinct retrieval stage, not the chapter's fusion policy.
Cited in 1 location
-
arXiv:2505.07897
Rando, S., et al. (2025), "LongCodeBench: Evaluating Coding LLMs at 1M Context Windows," arXiv:2505.07897.
Cited in 1 location
-
arXiv:2409.20489
Reid and colleagues (2024) provide directional evidence from a budget-constrained contextual-bandit formulation, theoretical analysis, and evaluations on real-world datasets.
Cited in 1 location
-
arXiv:1803.09578
Reimers & Gurevych (2018). Why Comparing Single Performance Scores Does Not Allow to Draw Conclusions About Machine Learning Approaches. arXiv:1803.09578. The 26% result.
Cited in 1 location
-
arXiv:2411.00998
Rosbach, E., Ganz, J., Ammeling, J., Riener, A., Aubreville, M. (2024). Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology. arXiv:2411.00998. This is the basis for the time-pressure paragraph. The setting is computational pathology, not software review.
Cited in 2 locations
-
arXiv:2605.15132
Rose, Evan, et al. 2026. APWA: A Distributed Architecture for Parallelizable Agentic Workflows. arXiv:2605.15132. Directional orchestration example.
Cited in 1 location
-
arXiv:2401.03729
Salinas and Morstatter (2024), arXiv:2401.03729. Trivial prompt edits changed more than 10 percent of predictions on some tested tasks.
Cited in 2 locations
-
arXiv:2208.09727
Sandoval and colleagues (2022/2023) conducted a controlled N=58 study on a low-level C linked-list task.
Cited in 1 location
-
arXiv:2208.06213
Sarkar, A., et al. (2022). What is it like to program with artificial intelligence? PPIG 2022. arXiv:2208.06213. This is the basis for the companion hypothesis that verification is the dominant cost. The magnitudes are not quantified.
Cited in 2 locations
-
arXiv:2410.18124
Saxena and colleagues (2024) extend that direction by treating availability as the objective, retaining at least two snapshots, and accounting for detection latency.
Cited in 1 location
-
arXiv:2310.11324
Sclar et al. (2023), FormatSpread, arXiv:2310.11324. Meaning-preserving prompt-format changes produced a median 7.5-point accuracy spread across the tested tasks, with larger task-level spreads.
Cited in 2 locations
-
arXiv:2605.03117
Seddik and colleagues (2026) provide strong evidence from controlled ablations of ARISE on SWE-bench Lite.
Cited in 1 location
-
arXiv:2604.05481
Sepidband, M., Viet Pham, H., Hemmati, H. (2026), "On the Role of Fault Localization Context for LLM-Based Program Repair," arXiv:2604.05481.
Cited in 1 location
-
arXiv:2604.08206
Shang, Wenlong. 2026. "Theater of Mind" for LLMs: A Cognitive Architecture Based on Global Workspace Theory. arXiv:2604.08206. Directional blackboard and global-workspace design; single-author architecture proposal, not a controlled measurement.
Cited in 2 locations
-
arXiv:2605.10913
Shepherd runtime substrate (Yu et al. 2026, arXiv:2605.10913); companion material named in the same synthesis without a paper identifier: OpenTelemetry GenAI conventions (CNCF 2025).
Cited in 1 location
-
arXiv:2502.17560
Singer and colleagues (2025) later report a shared multi-mission toolkit.
Cited in 1 location
-
arXiv:2605.00827
Singh Parmar (2026) reports one production Kubernetes synchronization case containing 67 steps.
Cited in 1 location
-
arXiv:2406.12624
Singh Thakur, Choudhary, Ramayapally, Vaidyanathan & Hupkes (2024), "Judging the Judges," arXiv:2406.12624. Chance-corrected alignment metrics, systematic leniency, and position effects.
Cited in 1 location
-
arXiv:2209.13085
Skalse, Howe, Krasheninnikov & Krueger (2022). Defining and Characterizing Reward Hacking. NeurIPS 2022. arXiv:2209.13085. Chapter 6, layer-signals-beyond-single-proxy.
Cited in 2 locations
-
arXiv:2605.24117
SkillEvolBench (Lei et al. 2026, arXiv:2605.24117), through the agentic-memory source synthesis. The underlying study measures no-skill and raw-trajectory controls in its setting; coverage matching, one-factor sweeps, and the writer information-flow constraint are transfers within the broader protocol.
Cited in 4 locations
-
arXiv:2502.03261
Somerstep, S., et al. (2025), "CARROT: A Cost Aware Rate Optimal Router," arXiv:2502.03261. Static model pool assumed.
Cited in 1 location
-
arXiv:2404.04059
Sterz, S., et al. (2024). On the Quest for Effectiveness in Human Oversight: Interdisciplinary Perspectives. ACM FAccT 2024. arXiv:2404.04059. Supports four conditions for effective oversight persons; gates failing any condition are compliance theater. The framework is motivated by EU AI Act Article 14, not…
Cited in 1 location
-
arXiv:2411.17501
Stroebl and colleagues (2024) showed that an imperfect verifier creates an accuracy ceiling that additional resampling cannot break.
Cited in 1 location
-
arXiv:2607.22368
Strong evidence for audited protocol defects and paired score inflation: Shao, J., Chen, H., Zhang, W., Pan, M., and Luo, B. (2026), "Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI," arXiv:2607.22368. The reported rates apply to the audited protocols, not all agent benchmarks.
Cited in 1 location
-
arXiv:2607.28587
Strong evidence for benchmark-specific PR-issue misalignment and detector accuracy: Wang, M., Xu, J., and He, P. (2026), "PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks," arXiv:2607.28587.
Cited in 1 location
-
arXiv:2607.25152
Strong evidence for externally grounded gating in the measured testbed: Park, H., and Choi, B. (2026), "When Do Agent Loops Mistake Stagnation for Progress?" arXiv:2607.25152. The 54-cycle result does not estimate a general agent failure rate.
Cited in 1 location
-
arXiv:2608.04804
Strong evidence for one benchmark-bound scout-and-fixer configuration and its no-router ablation: Bhola, I., Krishnan, A., and NS, M. (2026), "Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First," arXiv:2608.04804. The handoff, rather than the router, carried the measured result.
Cited in 1 location
-
arXiv:2608.00101
Strong evidence for production workload characterization, directional for capacity-policy transfer: Liu, B., et al. (2026), "Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale," arXiv:2608.00101.
Cited in 1 location
-
arXiv:2607.24882
Strong evidence for the benchmark measurements and controlled seed pilot: Qin, B., and Xie, Y. (2026), "Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents," arXiv:2607.24882. The benchmark supports stage-level evaluation rather than a universal retrieval-family ranking.
Cited in 1 location
-
arXiv:1608.07617
Strong evidence for the cheap-baseline comparison only: SWAY, Chen, J., et al. (2016), "Sampling as a Baseline Optimizer for Search-Based Software Engineering," arXiv:1608.07617. The study supports requiring a candidate optimizer to beat a cheap baseline; fixed-arrival replay is a transfer beyond its experiment.
Cited in 2 locations
-
arXiv:2605.12978
Strong evidence for the degradation mechanism: "Useful Memories Become Faulty When Continuously Updated by LLMs" (Zhang 2026), arXiv:2605.12978. The study measures degradation under repeated LLM memory updates; it does not compare immutable-source and rebuildable-distillate architectures.
Cited in 1 location
-
arXiv:2510.00615
Strong evidence for the measured failure-driven compression comparison: Kang, M., et al. (2025), "ACON: Optimizing Context Compression for Long-horizon LLM Agents," arXiv:2510.00615 (Microsoft). The study tests failure-guided compression on app, office, and question-answering benchmarks; transfer to coding agents…
Cited in 1 location
-
arXiv:2607.22520
Strong evidence for the narrower transition analysis: Tank, D., and Nama, B. (2026), "The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents," arXiv:2607.22520. Nearly 6,000 runs across two office-automation benchmarks and three model-harness stacks; transfer to coding-agent skill libraries is directional.
Cited in 1 location
-
arXiv:1909.09436
Strong evidence for the vocabulary-gap claim: CodeSearchNet (Husain 2019), arXiv:1909.09436. The study establishes the mismatch between natural-language queries and code vocabulary. It does not test the chapter's rank-fusion protocol; Yang et al. supply the direct comparison of combined lexical and semantic retrieval.
Cited in 1 location
-
arXiv:2607.11046
Strong evidence from file-level localization experiments by Caumartin and colleagues (2026) found that role-aware summaries outperformed file-path representations by up to 40% Hit@5 while using a representation footprint 10.4-20.9x smaller than raw source.
Cited in 1 location
-
arXiv:2605.08717
Strong evidence from Zhao and colleagues (2026) measured the gap across 257 cases: top-1 diagnosis accuracy was 65.37%, while recovery was 21.79%.
Cited in 1 location
-
arXiv:2607.27294
Strong evidence within the tested benchmark and configurations: Zhou, J., et al. (2026), "AgentS4D: Benchmarking Runtime Risks across the Execution Lifecycle of LLM-Based Workspace Agents," arXiv:2607.27294. Across 6,560 runs, 66.22 percent were both unsafe under a prespecified signal and complete; the result is not…
Cited in 1 location
-
arXiv:2503.11926
Strong evidence: Baker, B., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. OpenAI. arXiv:2503.11926.
Cited in 1 location
-
arXiv:2210.10760
Strong evidence: Gao, Schulman & Hilton (2022). Scaling Laws for Reward Model Overoptimization. ICML 2023. arXiv:2210.10760.
Cited in 1 location
-
DOI:10.1145/3656341
Sun, W., et al. (2024), "A Survey of Source Code Search: A 3-Dimensional Perspective," ACM Transactions on Software Engineering and Methodology, DOI:10.1145/3656341. The survey separates query, code, and matching components; transfer to repository agents remains directional.
Cited in 1 location
-
arXiv:2310.06770
SWE-bench (arXiv:2310.06770), from Jimenez et al. 2023. No figure carried.
Cited in 1 location
-
arXiv:2501.00539
Szeider (2024) offers a more recent prototype in which a language model edits declared constraint models while deterministic solvers consume them, although its authors characterize the experiments as non-rigorous benchmarks.
Cited in 1 location
-
arXiv:2605.29442
Tang et al., "How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions", arXiv:2605.29442, 2026.
Cited in 1 location
-
arXiv:2405.16081
Tang, N., Chen, M., Ning, Z., Bansal, A., Huang, Y., McMillan, C., Li, T. J.-J. (2024). A Study on Developer Behaviors for Validating and Repairing LLM-Generated Code Using Eye Tracking and IDE Actions. arXiv:2405.16081. Supports improved repair performance and greater verification effort with provenance disclosure;…
Cited in 1 location
-
arXiv:2508.02611
Tawosi and colleagues (2025) report directional evidence that summary-first reasoning can reduce repository context while retaining useful file and function localization.
Cited in 1 location
-
arXiv:2408.05861
Temporal KG memory (Kim et al.), arXiv:2408.05861, carried by the companion record on modeling time explicitly. Supports the two-time-axes aside only.
Cited in 2 locations
-
arXiv:2506.20430
The evidence is one directional explorer synthesis spanning the long-horizon engineering system described by Chen (2026), DeepRare from Zhao and colleagues (2026), and MemReader from Kang (2026).
Cited in 1 location
-
arXiv:2604.07877
The evidence is one directional explorer synthesis spanning the long-horizon engineering system described by Chen (2026), DeepRare from Zhao and colleagues (2026), and MemReader from Kang (2026).
Cited in 1 location
-
DOI:10.1109/TSE.2022.3165938
The lane also follows the more restrictive point made by Kitchenham and colleagues (2022): a mutable social-media post is not treated as a primary study merely because it is informative.
Cited in 1 location
-
arXiv:2605.17480
The tested multi-agent arrangements were more vulnerable in most configurations, and the Capability Paradox described by Liu (2026) directionally indicates that stronger worker models can make weaponized instructions more effective through persuasive linguistic certainty.
Cited in 1 location
-
DOI:10.1016/j.infsof.2018.09.006
This practitioner lane makes the review multivocal in the software-engineering sense described by Garousi, Felderer, and Mäntylä (2019): it combines scholarly and grey literature because operational mechanisms and incidents are often documented outside venues.
Cited in 1 location
-
arXiv:2512.20660
Thompson (2025) reports directional results across 13 models of 1.3B-15B parameters and three diagnostic probes, including up to 66-percentage-point reliability gains at 1.2-2.1x cost.
Cited in 1 location
-
arXiv:2509.23537
Tian, et al. 2025. "Beyond the Strongest LLM." arXiv:2509.23537. Directional evidence on multi-agent teams.
Cited in 2 locations
-
arXiv:2606.04056
Token Budgets (Khan 2026, arXiv:2606.04056). Together, the four synthesis items support only the direction of the checkpoint-placement claim. None is strong.
Cited in 2 locations
-
arXiv:1907.06250
Trofimov, Kuralenok, Marshalkin & Novikov (2019). Delivery, consistency, and determinism: rethinking guarantees in distributed stream processing. arXiv:1907.06250.
Cited in 1 location
-
DOI:10.1109/TSE.2023.3348172
Tufano, M., et al. (2024). "Code Review Automation: Strengths and Weaknesses of the State of the Art." IEEE Transactions on Software Engineering. DOI:10.1109/TSE.2023.3348172. Manual analysis of 2,291 predictions exposes capability boundaries and dataset defects; it does not evaluate this chapter's interface design.
Cited in 1 location
-
arXiv:2212.06823
Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M., Krishna, R. (2022/2023). Explanations Can Reduce Overreliance on AI Systems During Decision-Making. PACM HCI 7(CSCW1). arXiv:2212.06823. Five experiments treat overreliance as a cost-benefit choice and find greater engagement when…
Cited in 1 location
-
arXiv:2603.09619
Vishnyakova, O. (2026), arXiv:2603.09619. (Names the five context-quality criteria; position paper.)
Cited in 1 location
-
arXiv:2404.06203
Vogel et al. (2024), "A Comprehensive Benchmarking Analysis of Fault Recovery in Stream Processing Frameworks," arXiv:2404.06203 (JSS line).
Cited in 1 location
-
arXiv:2604.07667
Wang (2026), "Conformal Social Choice," arXiv:2604.07667.
Cited in 1 location
-
arXiv:2605.02431
Wei’s ARIADNE work (2026) presents a blackboard-driven search with active writes and a deadlock-breaking diversity check.
Cited in 1 location
-
arXiv:2605.14478
Weng, H., et al. (2026), "When Retrieval Hurts Code Completion: A Diagnostic Study of Stale Repository Context," arXiv:2605.14478.
Cited in 1 location
-
arXiv:2407.01489
Xia, C. S., et al. (2024), "Agentless: Demystifying LLM-based Software Engineering Agents," arXiv:2407.01489.
Cited in 1 location
-
arXiv:2408.01703
Xie, L., Zheng, C., Xia, H., Qu, H., Zhu-Tian, C. (2024). WaitGPT: Monitoring and Steering Conversational LLM Agent in Data Analysis with On-the-Fly Code Visualization. UIST 2024. arXiv:2408.01703. This prototype-scale single item is the basis for the live-state proposal and the thinnest support in the chapter.…
Cited in 2 locations
-
arXiv:2405.15793
Yang and colleagues (2024) also found that adding iterative search could perform worse than providing no search interface, showing how undirected exploration can consume budget without improving resolution.
Cited in 1 location
-
arXiv:2503.21710
Yang, B., et al. (2025), "Enhancing Repository-Level Software Repair via Repository-Aware Knowledge Graphs" (KGCompass), arXiv:2503.21710. (The 69.7 percent multi-hop figure is carried from v1; later versions report different headline values, so the inline link pins v1.)
Cited in 1 location
-
arXiv:2605.23929
Yang, Ya-Ting and Zhu (2026) directionally motivate allocating more budget to slowly saturating sequential stages, but their result is theory plus numerical illustration and assumes independent failures.
Cited in 1 location
-
arXiv:2507.18515
Yang, Z., et al. (2025), "A Deep Dive into Retrieval-Augmented Generation for Code Completion: Experience on WeChat," arXiv:2507.18515 (industrial study).
Cited in 1 location
-
arXiv:1909.13694
Ye, Martinez & Monperrus (2019), arXiv:1909.13694. [Strong evidence]
Cited in 1 location
-
arXiv:2502.00350
Yu and colleagues (2025) provide directional evidence from OrcaLoca on SWE-bench Lite.
Cited in 1 location
-
arXiv:2603.00520
Yu, B. et al. (2026), arXiv:2603.00520, SWE-ABS. [Strong evidence]
Cited in 1 location
-
arXiv:2503.07675
Yu, Junwei, Yepeng Ding, and Hiroyuki Sato. 2025. DynTaskMAS. arXiv:2503.07675. Directional parallel-versus-serial result on a seven-agent travel-planning workload; no controlled dynamic-versus-fixed comparison.
Cited in 1 location
-
arXiv:2506.09289
Yu, S. et al. (2025), arXiv:2506.09289, UTBoost. [Strong evidence]
Cited in 1 location
-
arXiv:2001.02114
Zhang and colleagues (2020) found in controlled experiments that confidence displays changed reliance in the expected direction, while joint accuracy improved only when people contributed complementary knowledge.
Cited in 1 location
-
arXiv:2405.00332
Zhang, H. et al. (2024), GSM1k, arXiv:2405.00332, Scale AI. [Strong evidence]
Cited in 1 location
-
arXiv:2505.23419
Zhang, L. et al. (2025), SWE-bench Goes Live!, arXiv:2505.23419, Microsoft Research. [Directional evidence]
Cited in 1 location
-
arXiv:2505.00212
Zhang, S., et al. (2025), "Which Agent Causes Task Failures and When? On Automated Failure Attribution of LLM Multi-Agent Systems," ICML 2025, arXiv:2505.00212. (Who&When: expert-annotated failure logs from 127 multi-agent systems; 53.5% agent-level and 14.2% step-level for the best automated method; some methods…
Cited in 2 locations
-
arXiv:2506.15655
Zhang, Y., et al. (2025), "cAST: Enhancing Code Retrieval-Augmented Generation with Structural Chunking via Abstract Syntax Tree," arXiv:2506.15655.
Cited in 1 location
-
arXiv:2605.26178
Zhao, et al. 2026. ATOM. arXiv:2605.26178. Audit-retained strong evidence on difficulty-conditioned topology across six short-form benchmarks.
Cited in 1 location
-
arXiv:2306.05685
Zheng et al. (2023), "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena," arXiv:2306.05685. Judge position, verbosity, and self-preference bias; inter-rater agreement against human labels.
Cited in 1 location
-
arXiv:2507.02825
Zhu and colleagues (2025) provide directional evidence through an Agentic Benchmark Checklist synthesized from surveyed failures.
Cited in 1 location
-
arXiv:2509.25370
Zhu and colleagues (2025) reported that AgentDebug’s root-cause isolation and targeted feedback outperformed retry-from-scratch baselines on trajectories from ALFWorld, GAIA, and WebShop.
Cited in 1 location
-
arXiv:2406.12045
τ-bench (Yao, Shinn, Razavi & Narasimhan 2024, arXiv:2406.12045). Introduces and measures \(\mathrm{pass}^{k}\) on interactive, multi-turn tool-use tasks with executable oracles. Using a team's own tasks and replaying the set per release are transfers beyond the measured findings.
Cited in 1 location