Research
Library
The collections I read from: ADS paper libraries I curate on NASA ADS, 10 thematic literature explorers that map recent work into themes, and a feed of what's been added lately. The sources; what I publish from them lives in the Digest.
Research libraries
NASA ADS ↗-
Coding Agents
92 documents
Software-engineering agents: architectures, multi-agent coding, and how developers work with them.
- Mining Architectural Quality Under Agentic AI Adoption: A Causal Study of Java Repositories
- Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents
-
Benchmarks
66 documents
Evaluating coding agents and code models on real software work.
- Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
- SWE-Explore: Benchmarking How Coding Agents Explore Repositories
-
Agentic Information Retrieval
62 documents
How LLM agents find information: dense retrieval and RAG foundations, reasoning-intensive and test-time-compute retrieval and reranking, agentic search loops, and code retrieval for coding agents.
- CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding
- SWE-Explore: Benchmarking How Coding Agents Explore Repositories
-
Agent Memory
158 documents
Long-horizon memory for LLM agents: storage, consolidation, and forgetting.
- Learning What to Remember: A Cognitively Grounded Multi-Factor Value Model for Agentic Memory
- G-Long: Graph-Enhanced Memory Management for Efficient Long-Term Dialogue Agents
-
Scientific Search & SciX
84 documents
Navigating scientific literature: NASA ADS / SciX information systems, scientific language models, and fine-grained classification of research text.
- Decades of Transformation: Evolution of the NASA Astrophysics Data System's Infrastructure
- Improving astroBERT Using Semantic Textual Similarity
-
Multi-Agent Orchestration
95 documents
Coordinating multiple LLM agents: controllers, shared context, task decomposition, and communication topologies.
- Reward Modeling for Multi-Agent Orchestration
- Decentralized Multi-Agent Systems with Shared Context
-
Code Retrieval & Enterprise Codebases
86 documents
Finding and navigating code at repository and enterprise scale for coding agents.
- SWE-Explore: Benchmarking How Coding Agents Explore Repositories
- CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding
Thematic explorers
Navigable, themed maps of recent literature. Each one starts from an ADS library and structures the papers into themes using SciX MCP.
-
Agentic Memory Systems
151 papers · 9 themes · 11 episodes
Procedural & SkillsReflection & ExperienceBenchmarksEval MethodologySynthetic DataArchitecturesSecurity & GovernanceApplications & PersonalizationForgetting & Consolidation -
Memory Design Considerations
82 papers · 11 themes · 1 episode
Retrieval & RankingConsolidation & DistillationKnowledge RepresentationTemporality & UpdatingForgetting & LifecycleStorage SubstrateMulti-Agent & Shared MemoryWorking Memory & ContextEvaluation & CostInterop, Schema & GovernanceFoundations & Landscape -
Enterprise Multi-Agent Reliability
144 papers · 8 themes
Reliability & failure modesRecovery & durable stateObservability & tracingEvaluation & assuranceCost, routing & schedulingTopology & coordinationSecurity & governanceHuman oversight & collaboration -
Multi-Agent Orchestration
127 papers · 5 themes
Foundations & TopologiesPatterns & FrameworksMemory in MASEnterprise & ProductionFrontier & Open Problems -
Code Retrieval & Enterprise Codebases
99 papers · 5 themes
Why Code != Text IRTechniques: Lexical/Neural/GraphRepository-Scale & Code GraphsEnterprise Codebase ChallengesFrontier & Open Problems -
Agentic Information Retrieval
30 papers · 4 themes · 2 episodes
How LLM agents find information — from dense retrieval and RAG to reasoning-intensive retrieval and test-time compute for ranking.
Dense retrieval & RAG foundationsAgentic search loopsReasoning-intensive retrievalTest-time compute for ranking -
Formal Specifications & Coding Agents
168 papers · 9 themes
Formal specification as the bottleneck in agent-written software: who writes the spec, how its quality is measured, and what a verifier can and cannot certify.
Specification foundationsWriting the specificationInvariants and proof obligationsRequirements to formal statementsVerification-aware code generationProof automation and prover agentsBenchmarks for specification and proof qualityCorrectness beyond proofsOpen questions and limits -
Physics AI & Causal World Models
80 papers · 8 themes · 8 episodes
The question I'm interested in is not whether learned models can predict physical systems; that result is increasingly established. It's what additional evidence is required before prediction can be interpreted as learned physical structure, causal understanding, or a model safe to use for counterfactual reasoning. I assembled this literature map to identify where those claims currently separate, which experiments discriminate between them, and which experiments still haven't been run.
Observations, assimilation & the sensor boundaryLearned weather & Earth-system predictionFoundation models & learned physical operatorsEvaluation, uncertainty & physical validityCausal structure & mechanistic understandingForecast to decisionWorld models, planning & physical controlIntervention & the boundary of evidence -
Where Business Meaning Lives
107 papers · 10 themes · 6 episodes
When an agent answers a question about a company's data, the schema tells it where the bytes are and nothing about what the business means by revenue, by churn, or by an active account. This map follows what happens to that missing meaning when it is written down as documentation, given a representation the model targets instead of SQL, compiled into a layer the model cannot bypass, or left for the model to infer while it writes the query. The sharpest measurements show that enforcement converts confident wrong answers into refusals more than it produces additional correct ones. I built this to work out which of those trades is worth making, where the studies disagree and why, and which experiments would settle the rest.
The schema underdetermines the questionSupplying meaning: what documentation buysRepresentation: the layer between language and SQLEnforcement: constraint instead of adviceThe boundary: wrong, uncertain, or unwillingThe harness became part of the systemHow text-to-SQL grades itselfWhat should generalize?Improving the system without gaming the benchmarkOriginal evidence: the Omni × LiveSQLBench benchmark -
The 30 Papers
27 papers · 7 themes
A beginner-shaped route through the 27 readings currently published at 30papers.com, from neural-network foundations to compression and complexity.
Build the working vocabularyDepth becomes trainableSequences meet the real worldAttention removes the bottleneckMemory becomes an interfaceScale becomes predictableCompression frames the whole field
Recently added
Recent papers across all libraries, ordered by publication date — the recency view of the corpus, from the hybrid scorer in my code-intelligence-digest app. For the written + audio issues drawn from this reading, see the Digest.
Knowledge linkage
Read the daily, weekly & selected digests → drawn from these collections. Concepts graph →