Research

Library

The collections I read from: ADS paper libraries I curate on NASA ADS, 10 thematic literature explorers that map recent work into themes, and a feed of what's been added lately. The sources; what I publish from them lives in the Digest.

Research libraries

NASA ADS ↗
  • Coding Agents

    92 documents

    Software-engineering agents: architectures, multi-agent coding, and how developers work with them.

    • Mining Architectural Quality Under Agentic AI Adoption: A Causal Study of Java Repositories
    • Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents

    Browse →

  • Benchmarks

    66 documents

    Evaluating coding agents and code models on real software work.

    • Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
    • SWE-Explore: Benchmarking How Coding Agents Explore Repositories

    Browse →

  • Agentic Information Retrieval

    62 documents

    How LLM agents find information: dense retrieval and RAG foundations, reasoning-intensive and test-time-compute retrieval and reranking, agentic search loops, and code retrieval for coding agents.

    • CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding
    • SWE-Explore: Benchmarking How Coding Agents Explore Repositories

    Browse →

  • Agent Memory

    158 documents

    Long-horizon memory for LLM agents: storage, consolidation, and forgetting.

    • Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory
    • G-Long: Graph-Enhanced Memory Management for Efficient Long-Term Dialogue Agents

    Browse →

  • Scientific Search & SciX

    84 documents

    Navigating scientific literature: NASA ADS / SciX information systems, scientific language models, and fine-grained classification of research text.

    • Decades of Transformation: Evolution of the NASA Astrophysics Data System's Infrastructure
    • Improving astroBERT Using Semantic Textual Similarity

    Browse →

  • Multi-Agent Orchestration

    95 documents

    Coordinating multiple LLM agents: controllers, shared context, task decomposition, and communication topologies.

    • Reward Modeling for Multi-Agent Orchestration
    • Decentralized Multi-Agent Systems with Shared Context

    Browse →

  • Code Retrieval & Enterprise Codebases

    86 documents

    Finding and navigating code at repository and enterprise scale for coding agents.

    • SWE-Explore: Benchmarking How Coding Agents Explore Repositories
    • CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding

    Browse →

Thematic explorers

Navigable, themed maps of recent literature. Each one starts from an ADS library and structures the papers into themes using SciX MCP.

  • Agentic Memory Systems

    151 papers · 9 themes · 11 episodes

    Procedural & SkillsReflection & ExperienceBenchmarksEval MethodologySynthetic DataArchitecturesSecurity & GovernanceApplications & PersonalizationForgetting & Consolidation

    Open explorer →

  • Memory Design Considerations

    82 papers · 11 themes · 1 episode

    Retrieval & RankingConsolidation & DistillationKnowledge RepresentationTemporality & UpdatingForgetting & LifecycleStorage SubstrateMulti-Agent & Shared MemoryWorking Memory & ContextEvaluation & CostInterop, Schema & GovernanceFoundations & Landscape

    Open explorer →

  • Enterprise Multi-Agent Reliability

    144 papers · 8 themes

    Reliability & failure modesRecovery & durable stateObservability & tracingEvaluation & assuranceCost, routing & schedulingTopology & coordinationSecurity & governanceHuman oversight & collaboration

    Open explorer →

  • Multi-Agent Orchestration

    127 papers · 5 themes

    Foundations & TopologiesPatterns & FrameworksMemory in MASEnterprise & ProductionFrontier & Open Problems

    Open explorer →

  • Code Retrieval & Enterprise Codebases

    99 papers · 5 themes

    Why Code != Text IRTechniques: Lexical/Neural/GraphRepository-Scale & Code GraphsEnterprise Codebase ChallengesFrontier & Open Problems

    Open explorer →

  • Agentic Information Retrieval

    30 papers · 4 themes · 2 episodes

    How LLM agents find information — from dense retrieval and RAG to reasoning-intensive retrieval and test-time compute for ranking.

    Dense retrieval & RAG foundationsAgentic search loopsReasoning-intensive retrievalTest-time compute for ranking

    Open explorer →

  • Formal Specifications & Coding Agents

    168 papers · 9 themes

    Formal specification as the bottleneck in agent-written software: who writes the spec, how its quality is measured, and what a verifier can and cannot certify.

    Specification foundationsWriting the specificationInvariants and proof obligationsRequirements to formal statementsVerification-aware code generationProof automation and prover agentsBenchmarks for specification and proof qualityCorrectness beyond proofsOpen questions and limits

    Open explorer →

  • Physics AI & Causal World Models

    80 papers · 8 themes · 8 episodes

    The question I'm interested in is not whether learned models can predict physical systems; that result is increasingly established. It's what additional evidence is required before prediction can be interpreted as learned physical structure, causal understanding, or a model safe to use for counterfactual reasoning. I assembled this literature map to identify where those claims currently separate, which experiments discriminate between them, and which experiments still haven't been run.

    Observations, assimilation & the sensor boundaryLearned weather & Earth-system predictionFoundation models & learned physical operatorsEvaluation, uncertainty & physical validityCausal structure & mechanistic understandingForecast to decisionWorld models, planning & physical controlIntervention & the boundary of evidence

    Open explorer →

  • Where Business Meaning Lives

    107 papers · 10 themes · 6 episodes

    When an agent answers a question about a company's data, the schema tells it where the bytes are and nothing about what the business means by revenue, by churn, or by an active account. This map follows what happens to that missing meaning when it is written down as documentation, given a representation the model targets instead of SQL, compiled into a layer the model cannot bypass, or left for the model to infer while it writes the query. The sharpest measurements show that enforcement converts confident wrong answers into refusals more than it produces additional correct ones. I built this to work out which of those trades is worth making, where the studies disagree and why, and which experiments would settle the rest.

    The schema underdetermines the questionSupplying meaning: what documentation buysRepresentation: the layer between language and SQLEnforcement: constraint instead of adviceThe boundary: wrong, uncertain, or unwillingThe harness became part of the systemHow text-to-SQL grades itselfWhat should generalize?Improving the system without gaming the benchmarkOriginal evidence: the Omni × LiveSQLBench benchmark

    Open explorer →

  • The 30 Papers

    27 papers · 7 themes

    A beginner-shaped route through the 27 readings currently published at 30papers.com, from neural-network foundations to compression and complexity.

    Build the working vocabularyDepth becomes trainableSequences meet the real worldAttention removes the bottleneckMemory becomes an interfaceScale becomes predictableCompression frames the whole field

    Open explorer →

Recently added

Recent papers across all libraries, ordered by publication date — the recency view of the corpus, from the hybrid scorer in my code-intelligence-digest app. For the written + audio issues drawn from this reading, see the Digest.

Paper Library Date Cites
Mining Architectural Quality Under Agentic AI Adoption: A Causal Study of Java Repositories Coding Agents 2026-06-00 0
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents Coding Agents 2026-06-00 0
The End of Code Review: Coding Agents Supersede Human Inspection Coding Agents 2026-06-00 0
Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull Requests Coding Agents 2026-06-00 0
Understanding the Rejection of Fixes Generated by Agentic Pull Requests -- Insights from the AIDev Dataset Coding Agents 2026-06-00 0
Decentralized Multi-Agent Systems with Shared Context Coding Agents 2026-06-00 0
Frontier Coding Agents Use Metaprogramming to Adapt to Unfamiliar Programming Languages Coding Agents 2026-06-00 0
Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests Benchmarks 2026-06-00 0
SWE-Explore: Benchmarking How Coding Agents Explore Repositories Benchmarks 2026-06-00 0
SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? Benchmarks 2026-06-00 0
CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding Benchmarks 2026-06-00 0
Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory Agent Memory 2026-06-00 0
G-Long: Graph-Enhanced Memory Management for Efficient Long-Term Dialogue Agents Agent Memory 2026-06-00 0
Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory Agent Memory 2026-06-00 0
Less Context, More Accuracy: A Bi-Temporal Memory Engine for LLM Agents Where a Lean Retrieved Context Beats the Full History Agent Memory 2026-06-00 0
Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents Agent Memory 2026-06-00 0
Reward Modeling for Multi-Agent Orchestration Multi-Agent Orchestration 2026-06-00 0
Verbal-R3: Verbal Reranker as the Missing Bridge between Retrieval and Reasoning Agentic Information Retrieval 2026-05-00 0
GRC: Unifying Reasoning-Driven Generation, Retrieval and Compression Agentic Information Retrieval 2026-05-00 0
RICE-PO: Turning Retrieval Interactions into Credit Signals for Reasoning Agents Agentic Information Retrieval 2026-05-00 0
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems Agent Memory 2026-05-00 0
The Time is Here for Just-in-Time Systems: Challenges and Opportunities Agent Memory 2026-05-00 0
Rethinking How to Remember: Beyond Atomic Facts in Lifelong LLM Agent Memory Agent Memory 2026-05-00 0
Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents Agent Memory 2026-05-00 0
Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems Agent Memory 2026-05-00 0
GRAVITY: Architecture-Agnostic Structured Anchoring for Long-Horizon Conversational Memory Agent Memory 2026-05-00 0
From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms Agent Memory 2026-05-00 0
STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? Agent Memory 2026-05-00 0
Hijacking Agent Memory: Stealthy Trojan Attacks Through Conversational Interaction Agent Memory 2026-05-00 0
MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts Agent Memory 2026-05-00 0

Knowledge linkage