Agent architecture
Structural design of agent systems: planning loops, sub-agent decomposition, state management, and control flow.
tagged in 1 of 82 digest issues, most recently 2026-06-22
Agent memory
How agents store, retrieve, and forget context across turns and sessions; memory architectures and their design tradeoffs.
tagged in 42 of 82 digest issues, most recently 2026-07-15
Digest issues
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- Claude Cowork moves to the cloud, and 90% of its sessions aren't coding — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- AI pull requests carry 1.7× more defects, and review didn't scale to match — 2026-06-27
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Anthropic Puts Claude in Slack — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Enhancing Developer Productivity with Google Colab CLI and Agentic Observability — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
Agent observability
Tracing, monitoring, and debugging agent runs in production: telemetry, replay, and failure analysis.
tagged in 1 of 82 digest issues, most recently 2026-06-25
Agent tooling
Harnesses, skills, and tool interfaces that let agents act: the plumbing between a model and the systems it operates.
tagged in 37 of 82 digest issues, most recently 2026-07-13
Digest issues
- Frontier agents ace one task and stall across a stream of them — 2026-07-13
- OpenAI claims a math proof, walks back its launch, and the model race turns to cost — 2026-07-12
- Claude Cowork moves to the cloud, and 90% of its sessions aren't coding — 2026-07-11
- GPT-5.6 lands, and the coding-model price floor drops again — 2026-07-10
- Grok 4.5 Ships as Cursor's First General-Purpose Model, and OpenAI Retracts SWE-Bench Pro — 2026-07-09
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- x402 payments reach both edge networks, and GPT-5.6 Sol Ultra heads to Codex — 2026-07-06
- Sonnet 5 Lands, Fable 5 Returns, and ZCode Undercuts Everyone on Price — 2026-07-06
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Fable 5 completes the task as a test and refuses it as production work — 2026-07-04
- The field can write more code than it can understand — 2026-07-03
- Claude Tag lands 65% of Anthropic's internal PRs while the World's Fair argues over the outer loop — 2026-07-02
- Claude Sonnet 5 lands everywhere at once — 2026-07-01
- The frontier model becomes the supervisor — 2026-06-30
- Open weights became the default coding model the week the frontier got restricted — 2026-06-29
- The floor is rising: open models, the coding-speed paradox, and the bill for the buildout — 2026-06-29
- GPT-5.6 and Mythos 5 ship behind a government access gate — 2026-06-28
- OpenAI ships GPT-5.6 behind a government gate, the same day Washington un-blocks Mythos 5 — 2026-06-27
- Codex usage at OpenAI jumped 56x, and the agent stack started grading itself — 2026-06-26
- Open weights match Opus at half the cost, and ship free in Devin — 2026-06-25
- Anthropic Puts Claude in Slack — 2026-06-24
- Cyber models ship faster than the rules for them — 2026-06-23
- GLM 5.2 edges past Sonnet on 1,000 coding tasks as Claude's error rates spike — 2026-06-22
- GLM-5.2 takes the open-model crown while agents settle into production — 2026-06-22
- A single page can RCE your agent's host — 2026-06-21
- Anthropic resets every usage limit while negotiating Fable back from a US ban — 2026-06-20
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- Two labs put their models on the lab bench — 2026-06-18
- GLM-5.2 cracks the open-weight coding frontier, and Cursor goes to SpaceX — 2026-06-17
- The Fable 5 "jailbreak" was "fix this code" — 2026-06-16
- When the loop closes, verification becomes the job — 2026-06-15
- Anthropic shipped the best coding model measured, then the government pulled it — 2026-06-15
- The model layer becomes a regulated surface — 2026-06-14
- The day the US government switched off Fable 5 — 2026-06-13
- OpenAI buys Ona while Fable 5 starts downgrading itself — 2026-06-12
- Fable 5 scores 91, real code scores 13 — 2026-06-11
- Claude Fable 5 arrives at twice Opus pricing, and Cognition's day-old FrontierCode crowns it #1 — 2026-06-10
Agentic coding
Agents that write, refactor, and maintain software: coding assistants, autonomous dev loops, and the workflows around them.
tagged in 51 of 82 digest issues, most recently 2026-07-15
Digest issues
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- Grok Build CLI uploaded whole repos to a Google bucket; Codex hit 7M users — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- Frontier agents ace one task and stall across a stream of them — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Sonnet 5 Lands, Fable 5 Returns, and ZCode Undercuts Everyone on Price — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The field can write more code than it can understand — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- The floor is rising: open models, the coding-speed paradox, and the bill for the buildout — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- AI pull requests carry 1.7× more defects, and review didn't scale to match — 2026-06-27
- The agent scaffold is now the contested layer — 2026-06-26
- Codex usage at OpenAI jumped 56x, and the agent stack started grading itself — 2026-06-26
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
AI agents
Systems that plan, call tools, and act over multiple steps to accomplish a goal.
tagged in 0 of 82 digest issues
AI and the labor market
AI's effect on jobs and work: displacement, augmentation, and workforce shifts.
tagged in 3 of 82 digest issues, most recently 2026-07-03
AI economics
Cost structure of AI: token pricing, inference margins, usage limits, and unit economics.
tagged in 13 of 82 digest issues, most recently 2026-07-12
AI for science
AI applied to scientific discovery and research workflows, from literature synthesis to hypothesis generation.
tagged in 7 of 82 digest issues, most recently 2026-07-14
AI governance
Rules for AI systems: regulation, policy, export controls, and organizational governance of agents.
tagged in 13 of 82 digest issues, most recently 2026-07-13
AI industry
The business landscape of AI: labs, funding, acquisitions, and competitive strategy.
tagged in 5 of 82 digest issues, most recently 2026-07-14
AI infrastructure
Compute, serving, and platform layers under AI systems: GPUs, inference stacks, and datacenter buildout.
tagged in 9 of 82 digest issues, most recently 2026-06-29
AI safety
Preventing harmful model and agent behavior: alignment, oversight, and safety evaluation.
tagged in 9 of 82 digest issues, most recently 2026-07-09
AI security
Securing AI systems and using AI in security: prompt injection, sandboxing, supply-chain risk, and agent attack surfaces.
tagged in 18 of 82 digest issues, most recently 2026-07-14 · 15 papers · 1 explorer section
Digest issues
- Grok Build CLI uploaded whole repos to a Google bucket; Codex hit 7M users — 2026-07-14
- OpenAI claims a math proof, walks back its launch, and the model race turns to cost — 2026-07-12
- Claude Cowork moves to the cloud, and 90% of its sessions aren't coding — 2026-07-11
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Fable 5 completes the task as a test and refuses it as production work — 2026-07-04
- The field can write more code than it can understand — 2026-07-03
- Open weights became the default coding model the week the frontier got restricted — 2026-06-29
- The floor is rising: open models, the coding-speed paradox, and the bill for the buildout — 2026-06-29
- Open weights match Opus at half the cost, and ship free in Devin — 2026-06-25
- Anthropic Puts Claude in Slack — 2026-06-24
- Cyber models ship faster than the rules for them — 2026-06-23
- A single page can RCE your agent's host — 2026-06-21
- Two labs put their models on the lab bench — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- When the loop closes, verification becomes the job — 2026-06-15
- The model layer becomes a regulated surface — 2026-06-14
- The day the US government switched off Fable 5 — 2026-06-13
Papers
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection — Attackers compromise LLM-integrated apps by planting instructions in content the model later retrieves (web pages, documents) — no direct access needed.
- AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents — A dynamic benchmark of realistic tool-using tasks for measuring prompt-injection attacks and defenses — and showing current defenses are far from complete.
- G-Safeguard: A Topology-Guided Security Lens and Treatment on LLM-based Multi-Agent Systems — Detects anomalies over the agent communication graph and intervenes to defend a multi-agent system against attacks that propagate between agents.
- Red-Teaming LLM Multi-Agent Systems via Communication Attacks — Injects malicious messages that propagate through inter-agent communication channels, compromising a multi-agent system from the inside.
- Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions — Surveys the security threat landscape of MCP — the tool-connection standard now powering most agentic systems — and its expanding attack surface.
- MPAC: A Multi-Principal Agent Coordination Protocol (extends MCP + A2A)
- Architecture Matters for Multi-Agent Security — The same task under different multi-agent architectures has different security properties.
- AgentShield: Deception-based Compromise Detection for Tool-using LLM Agents — Plants deception traps (fake tools/credentials) that detect indirect-prompt-injection compromises which slip past prevention — cross-lingual, near-zero false positives.
- AI Assurance: A Comprehensive Testing Strategy for Enterprise AI Systems
- Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents — PORTICO is a reference monitor that gives coding-agent tools a bounded capability lifetime: an explicit task contract compiles into an initial envelope plus grant, trusted-closure, and global-deny rules, and a request-grant-invoke protocol mints opaque, epoch-bound handles that are revoked when the subgoal closes, so authority cannot linger or be replayed once its justifying episode ends.
- AgentLens: Interpretable Safety Steering via Mechanistic Subspaces for Multi-Turn Coding Agent — A white-box runtime defense for multi-turn coding agents that drive a shell in Docker. At each step it reads the last-token hidden state at one layer and runs a single linear probe to flag a harmful execution state; on a flag it steers the representation inside a sparse 10-dimensional subspace, adapting only the steering strength via an LLM judge scoring safety (0.6) and utility (0.4). It detects current-step risk at 97.32% average accuracy, predicts the next harmful command at up to 96.77% lookahead, and cuts average attack success from 85.99% to 13.36% (72.63 pp), beating RepE and self-reminder. Negative steering of the same subspace flips genuine refusals into malicious commands (100% ASR on LLaMA), confirming the direction is causal not lexical; under prompt injection the probe misfires but applying the steering direction still drops ASR 86.7%->6.7%.
- ACRFence: Preventing Semantic Rollback Attacks in Agent Checkpoint-Restore
- AgentSight: System-Level Observability for AI Agents Using eBPF
- NIST AI Risk Management Framework (AI RMF 1.0)
- OWASP Top 10 for LLM Applications — A community-maintained catalogue of the top LLM-application security risks — prompt injection, insecure output handling, excessive agency, supply chain, and more.
Code intelligence
Understanding codebases at scale: search, navigation, and agents that reason over source.
tagged in 0 of 82 digest issues
Code review
Automated and agent-assisted review of code changes: quality gates, review bots, and human-agent review workflows.
tagged in 7 of 82 digest issues, most recently 2026-07-13
Computer use
Agents that operate GUIs and browsers directly: screen perception, action models, and their reliability limits.
tagged in 1 of 82 digest issues, most recently 2026-06-25
Context engineering
Deciding what goes into a model's context window and when: packing, pruning, and structuring working context for long-running tasks.
tagged in 4 of 82 digest issues, most recently 2026-07-12
Developer productivity
How AI tooling changes software work: measured impact, adoption patterns, and workflow shifts.
tagged in 5 of 82 digest issues, most recently 2026-07-10
Enterprise adoption
How organizations deploy AI in production: procurement, integration, and the gap between demos and durable systems.
tagged in 2 of 82 digest issues, most recently 2026-06-22
Evaluation
Measuring whether AI systems work: benchmark design, eval harnesses, and comparison of models, agents, and search systems.
tagged in 53 of 82 digest issues, most recently 2026-07-15 · 18 papers · 2 explorer sections
Digest issues
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- Grok 4.5 Ships as Cursor's First General-Purpose Model, and OpenAI Retracts SWE-Bench Pro — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Claude Tag lands 65% of Anthropic's internal PRs while the World's Fair argues over the outer loop — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- The frontier model becomes the supervisor — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- AI pull requests carry 1.7× more defects, and review didn't scale to match — 2026-06-27
- The agent scaffold is now the contested layer — 2026-06-26
- Codex usage at OpenAI jumped 56x, and the agent stack started grading itself — 2026-06-26
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- GLM 5.2 edges past Sonnet on 1,000 coding tasks as Claude's error rates spike — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Fable 5 scores 91, real code scores 13 — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- Claude Fable 5 arrives at twice Opus pricing, and Cognition's day-old FrontierCode crowns it #1 — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Enhancing Developer Productivity with Google Colab CLI and Agentic Observability — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
Papers
- Evaluating Very Long-Term Conversational Memory of LLM Agents — Introduces LoCoMo, the canonical long-term conversational-memory benchmark: very-long dyadic (person-person) dialogues spanning up to ~35 sessions / hundreds of turns, generated via a machine-human pipeline with personas and temporal event graphs, plus a QA test suite, event-summarization, and multimodal-dialogue-generation tasks.
- Recent Trends in Personalized Dialogue Generation: A Review of Datasets, Methodologies, and Evaluations — Surveys 22 personalized-dialogue datasets and 17 seminal works (2021-2023), cataloging persona/personalization datasets, problem types, and evaluation facets/metrics for personalized dialogue.
- Agent Workflow Memory — Introduces Agent Workflow Memory, inducing reusable workflows from past trajectories to improve long-horizon web-navigation agents, evaluated on agent-trajectory benchmarks (e.g. web tasks) rather than conversational QA.
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory — Defines LongMemEval, a chat-assistant memory benchmark organized around five core abilities (information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention), with controllable interaction histories so memory load can be scaled independently of the target evidence.
- REALTALK: A 21-Day Real-World Dataset for Long-Term Conversation — Releases REALTALK, 21 days of genuine human-human messaging conversations (not LLM-synthesized), used to study long-term memory probing and persona/emotional-intelligence simulation against the distribution gap between synthetic and real dialogue.
- Evaluating LLM-based Agents for Multi-Turn Conversations: A Survey — PRISMA-based survey of evaluation methods for LLM agents in multi-turn conversational settings, cataloging benchmarks, metrics, and methodologies across the multi-turn evaluation landscape.
- ImplicitMemBench: Measuring Unconscious Behavioral Adaptation in Large Language Models — Proposes ImplicitMemBench to measure implicit (procedural) memory, where past experience becomes automated behavior, rather than the explicit fact-recall that existing memory benchmarks test.
- Evaluating Memory Capability in Continuous Lifelog Scenario — Argues existing memory benchmarks cover only Person-AI and dyadic Person-Person dialogue and neglects 'continuous dialogue lifelogs' from always-on wearables; introduces lifelog memory evaluation resources (LifeMem / EgoMem first-person scenarios) for ambient continuously-recorded conversation.
- Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo — Pins down LoCoMo's exact composition (ACL 2024 version: 10 conversations, 1,813 questions) and argues every LoCoMo system treats memory as retrieval over dialogue segments while none reports adversarial robustness (refusing questions about facts never disclosed); proposes that abstention/hallucination-resistance metric as a missing axis.
- AgenticAI-DialogGen: Topic-Guided Conversation Generation for Fine-Tuning and Evaluating Short- and Long-Term Memories of LLMs — Presents a topic-guided multi-agent conversation-generation pipeline that synthesizes dialogues specifically for fine-tuning and evaluating both short- and long-term memory.
- GroupMemBench: Benchmarking LLM Agent Memory in Multi-Party Conversations — Introduces GroupMemBench, targeting memory in multi-party (group) conversations where the agent must attribute facts to specific speakers across many participants, a setting that dyadic benchmarks like LoCoMo/LongMemEval do not cover.
- MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models — Introduces MemLens to benchmark long-term memory in vision-language models across long multimodal interactions, evaluating whether memory methods preserve evidence needed for later recall.
- MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory — Proposes MemEye, a visual-centric evaluation framework that specifically tests whether agents preserve the visual evidence required for later recall in multimodal memory.
- MemFail: Stress-Testing Failure Modes of LLM Memory Systems — Builds a stress-test benchmark that systematically enumerates and probes failure modes of external memory systems (consistency breaks over long-horizon interactions) rather than aggregate accuracy.
- Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison — Constructs a diagnostic testbed for selective QA over conflicting multi-source personal memory, where the system must decide how to use contradictory stored facts, plus a comparison of methods on that testbed.
- GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge — Separates the value of a substrate from raw accuracy: per-backend cost is measured, and a similarity sweep shows memory pays only above a 'copyability threshold' — when the retrieved case is a near-duplicate (cosine ≳ 0.8) accuracy jumps +12 to +13.5 pp, while below it nothing helps; the gain is answer retrieval, not method transfer, and a 4.5× larger backbone steepens the near-duplicate step (+22.5 to +28.5 pp) yet still cannot extract a transferable method from a worked example.
- StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance — A streaming benchmark sourced from real EgoLife egocentric lifelogs rather than scripted or synthesized dialogue: each five-minute segment is mined for a hidden 'evidence anchor' (a user-specific preference, plan, or capability) that spawns a two-step task sequence — an initial task that needs the evidence and a follow-up task grounded in the same anchor; eight memory systems are evaluated across two backbones (DeepSeek-V4-Flash, Gemini-3-Flash).
- MemTrace: Probing What Final Accuracy Misses in Long-Term Memory — MemTrace contains 835 typed knowledge points drawn from 20 users, expanded into 15,422 question rows and over 200,000 scored answers, and is used to evaluate 13 memory-system configurations spanning four paradigms: long-context models, retrieval-augmented systems, external-memory stores, and agentic-memory architectures.
Generative UI
Interfaces generated on the fly by models: dynamic layouts, artifacts, and agent-driven frontends.
tagged in 1 of 82 digest issues, most recently 2026-07-03
Human in the loop
Where people sit in agent workflows: approval gates, escalation, and oversight interfaces.
tagged in 0 of 82 digest issues
Information retrieval
Finding the right information at the right time: ranking, hybrid search, and retrieval over scientific and code corpora.
tagged in 38 of 82 digest issues, most recently 2026-07-15 · 3 papers · 1 explorer section
Digest issues
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- Same pass rate, half the cost: coding-agent evals turn to cost per task — 2026-07-14
- The context bill, and where coding agents break — 2026-07-13
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Microsoft measured a 24% PR lift from CLI coding agents — 2026-07-04
- Reasoning effort buys reliability, testing tools buy cost — 2026-07-03
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- The agent scaffold is now the contested layer — 2026-06-26
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Enhancing Developer Productivity with Google Colab CLI and Agentic Observability — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
Papers
- Engram: A Bi-Temporal Memory Engine Where a Lean Retrieved Context Beats the Full History — The read path retrieves through four channels in parallel — dense semantic, BM25 lexical, graph n-hop from query entities, and recency/salience — fuses them with Reciprocal Rank Fusion, then applies an 'as-of' filter and an abstention gate; the assembled context is hybrid (conflict-resolved facts plus raw session chunks) because facts alone lose recall.
- Infini Memory: Maintainable Topic Documents for Long-Term LLM Agent Memory — At read time an agentic procedure lets the LLM iteratively call memory tools — inspect intermediate results, expand local context around matches, assemble evidence — rather than take a single top-k step; the agentic variant beats a hybrid summary+BM25 reader (79.3% vs 76.0% on LongMemEval_S).
- T-Mem: Memory That Anticipates, Not Archives — T-Mem retrieves through a top-down topic -> scene -> item cascade, scoring each layer with reciprocal rank fusion over both a shared BM25 lexical index and a per-type dense index (bge-m3). On top of these node indices, four write-time 'trigger' families surface host nodes on the query's behalf: items expose three independently-encoded views (concept-only, bridge-only, joint) and the trigger score is the nan-aware max across views, attributed back to the host item. Scenes/items reached through associative triggers bypass the topic prefilter, because gating them by surviving topics would re-impose the similarity-only neighbourhood the system is built to escape.
Memory consolidation
Distilling episodic traces into durable knowledge and skills, and deciding what an agent should forget.
18 papers · 2 explorer sections · tagged in 0 of 82 digest issues
Papers
- Lifelong Learning of Large Language Model based Agents: A Roadmap — Roadmap framing continual/incremental learning for LLM agents, organizing memory, forgetting, and knowledge-accumulation challenges.
- A-MEM: Agentic Memory for LLM Agents — Zettelkasten-style self-organizing agent memory with dynamic linking and note evolution as new memories arrive.
- Neuron-level Balance between Stability and Plasticity in Deep Reinforcement Learning — Per-neuron stability/plasticity balancing to address the stability-plasticity dilemma in deep RL agents.
- SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment — Benchmark moving beyond episodic task scoring to evaluate self-evolving agents across time, targeting 'episodic amnesia' and static-toolset limits.
- Learning to Forget -- Hierarchical Episodic Memory for Lifelong Robot Deployment — H2-EMV: LM-based relevance estimation drives selective forgetting under learned natural-language rules updated by user feedback; measures QA accuracy retention against memory-size and query-compute savings.
- Time is Not a Label: Continuous Phase Rotation for Temporal Knowledge Graphs and Agentic Memory — Encodes time as continuous phase rotation (not a discrete label) in temporal KG / agentic memory so validity intervals and obsolescence are representable.
- Cooperative Memory Paging with Keyword Bookmarks for Long-Horizon LLM Conversations — Evicts old content past the context window but indexes it with keyword bookmarks so the model can recover evicted memory on demand.
- Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments — Liquid-Glass-Crystal three-phase consolidation governed by an Itô SDE / Fokker-Planck (closed-form Beta stationary dist.); proves convergence + memory-capacity bounds and empirically reports forward transfer and forgetting reductions.
- SCM: Sleep-Consolidated Memory with Algorithmic Forgetting for Large Language Models — Memory architecture drawing on neuroscientific sleep-consolidation principles, pairing offline consolidation with an explicit algorithmic-forgetting operator to bound store growth.
- Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents — Typed semantic memory using information-theoretic retrieval scoring to govern what is retained/surfaced across sessions for long-horizon agents.
- Evolve: A Persistent Knowledge Lifecycle for Small Language Models — Teacher-compiled persistent knowledge store refined through sleep consolidation and usage-driven retention/decay for small local LMs.
- ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems — Replaces system-engineering memory metaphors (paging/flat stores) with a 7-layer neuroscience-grounded architecture incorporating consolidation and forgetting layers.
- STALE: Can LLM Agents Know When Their Memories Are No Longer Valid? — Benchmark of 400 expert-validated conflict scenarios / 1,200 queries (contexts to 150K tokens) probing belief revision over time via three dimensions: State Resolution, Premise Resistance, Implicit Policy Adaptation; introduces 'Implicit Conflict' failure mode and CUPMem write-time-revision baseline.
- EvolveMem: Self-Evolving Memory Architecture via AutoResearch for LLM Agents — Self-evolving memory that treats retrieval infrastructure as mutable, auto-researching its own organization across multi-session operation.
- NeuSymMS: A Hybrid Neuro-Symbolic Memory System for Persistent, Self-Curating LLM Agents — Neuro-symbolic, self-curating memory where symbolic structure governs update/curation of persistent cross-session user knowledge.
- Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents — Frames retention/eviction as a constrained multi-step stochastic optimization over a hard storage budget, with the per-step reward explicitly charging miss penalties, reacquisition delay, and stale-information use; the learned policy keeps a smaller, evidence-denser set rather than greedily filling the budget (LoCoMo budget 128: F1 0.302 at 0.76 occupancy vs Mixed-Score 0.069 at 0.99).
- TokenPilot: Cache-Efficient Context Management for LLM Agents — Lifecycle-Aware Eviction tracks each context segment through three states active -> completed -> evictable, gated on evidence that a sub-task achieved its objective and on residual utility. A completed segment is not immediately purged; it retains its physical cache slots as long as residual relevance to ongoing interactions is non-zero, and only an evictable segment (utility decayed to zero) is removed in a single-pass structural purge. State estimation runs via a lightweight zero-shot validator over a compressed historical view in conservative batches of B turns rather than every step.
- Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents — Frames consolidation on the complementary-learning-systems analogy (fast episodic plus slow consolidating stores) but states it is motivation only, not a biological model. Consolidation is measured by write economy and bounded drift, not task accuracy: EVAF reaches goal persistence and post-unload recovery of 0.812–0.904 with only 2–3 parametric writes per 200 events (L2 drift ~21–29), while Naive-LoRA writes every event (200 writes, L2 drift ~67 on TinyLlama, ~119 on GPT-2) and still fails the goal layer — so writing everything is not enough. Stability-plasticity is handled with replay plus an L2/EWC-style anchor. The unresolved boundary is stale-memory obsolescence: on public Memora event streams EVAF improves forgetting-absence only 91/222 to 95/222 (McNemar p=0.57, not significant), and negative-gradient/anti-training forgetting variants were unstable — append-only selective consolidation does not solve delete/update validity.
Model Context Protocol
The open protocol connecting models to tools and data sources, and the server/client ecosystem built on it.
tagged in 5 of 82 digest issues, most recently 2026-07-06
Model releases
New model launches and availability changes across frontier and open providers.
tagged in 26 of 82 digest issues, most recently 2026-07-14
Digest issues
- Grok Build CLI uploaded whole repos to a Google bucket; Codex hit 7M users — 2026-07-14
- OpenAI claims a math proof, walks back its launch, and the model race turns to cost — 2026-07-12
- GPT-5.6 lands, and the coding-model price floor drops again — 2026-07-10
- Grok 4.5 Ships as Cursor's First General-Purpose Model, and OpenAI Retracts SWE-Bench Pro — 2026-07-09
- GPT-5.6 gets a launch date, and an agent leaks a private repo — 2026-07-08
- x402 payments reach both edge networks, and GPT-5.6 Sol Ultra heads to Codex — 2026-07-06
- Sonnet 5 Lands, Fable 5 Returns, and ZCode Undercuts Everyone on Price — 2026-07-06
- Fable finds five release blockers in sqlite-utils, two days before the price cliff — 2026-07-05
- Fable 5 completes the task as a test and refuses it as production work — 2026-07-04
- Claude Sonnet 5 lands everywhere at once — 2026-07-01
- The frontier model becomes the supervisor — 2026-06-30
- Open weights became the default coding model the week the frontier got restricted — 2026-06-29
- GPT-5.6 and Mythos 5 ship behind a government access gate — 2026-06-28
- OpenAI ships GPT-5.6 behind a government gate, the same day Washington un-blocks Mythos 5 — 2026-06-27
- Cyber models ship faster than the rules for them — 2026-06-23
- GLM-5.2 takes the open-model crown while agents settle into production — 2026-06-22
- Anthropic resets every usage limit while negotiating Fable back from a US ban — 2026-06-20
- GLM-5.2 passes the vibe check, and the agent-safety papers pile up — 2026-06-19
- Two labs put their models on the lab bench — 2026-06-18
- GLM-5.2 cracks the open-weight coding frontier, and Cursor goes to SpaceX — 2026-06-17
- Anthropic shipped the best coding model measured, then the government pulled it — 2026-06-15
- The model layer becomes a regulated surface — 2026-06-14
- The day the US government switched off Fable 5 — 2026-06-13
- OpenAI buys Ona while Fable 5 starts downgrading itself — 2026-06-12
- Fable 5 scores 91, real code scores 13 — 2026-06-11
- Claude Fable 5 arrives at twice Opus pricing, and Cognition's day-old FrontierCode crowns it #1 — 2026-06-10
Multi-agent orchestration
Coordinating multiple agents on shared work: topologies, delegation patterns, shared memory, and production reliability.
tagged in 40 of 82 digest issues, most recently 2026-07-15
Digest issues
- Agent skills encode preconditions, and agents violate them up to 70% of the time — 2026-07-15
- The harness, not the model, moved this week's numbers — 2026-07-13
- The harness moved the bill more than the model did — 2026-07-12
- The harness, not the model, is where the reviews got cheaper — 2026-07-11
- A Third of SWE-Bench Pro's Grades Don't Survive a Second Look — 2026-07-10
- A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget — 2026-07-09
- The harness is where the leverage went — 2026-07-08
- Clean code and memory buy coding agents efficiency, not success — 2026-07-06
- Delegation Is Not Management — 2026-07-06
- Agent cost numbers are wrong until you count the whole tree — 2026-07-05
- The Scores Fall Apart When You Change the Machine — 2026-07-02
- Claude Tag lands 65% of Anthropic's internal PRs while the World's Fair argues over the outer loop — 2026-07-02
- Agent memory grows a stack, and an attack surface — 2026-07-01
- When the checker becomes the target — 2026-06-30
- Agent risk lives in the repository, not the agent — 2026-06-29
- Submit 100%, Resolve 44%: Agent Evals Move Past Completion — 2026-06-29
- Verification has a price, and agents are overpaying — 2026-06-28
- The agent scaffold is now the contested layer — 2026-06-26
- Coding agents keep declaring victory they didn't earn — 2026-06-25
- Measuring How Agents Work, Not Just Whether They Finished — 2026-06-24
- Submit rate isn't resolve rate — 2026-06-23
- Open models caught Sonnet on coding tasks; the harness is where the gap moved — 2026-06-22
- The harness moves agent scores as much as the model does — 2026-06-22
- The bottleneck moved to the scaffolding — 2026-06-21
- Coding-agent reliability gets attacked from the harness and the crowd — 2026-06-20
- Coding agents crack at turn five; the day's work is in the harness layer — 2026-06-19
- We measure the model; the harness decides the run — 2026-06-18
- The bottleneck in agent memory is evidence use, not retrieval — 2026-06-17
- Memory has a bill, and the field started reading it — 2026-06-16
- Memory only helps your agent when it's looking at a near-duplicate — 2026-06-15
- When the loop closes, verification becomes the job — 2026-06-15
- The benchmarks caught up to the slop — 2026-06-15
- Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip — 2026-06-14
- Same model, +10 points on Oolong: today's gains came from the harness — 2026-06-13
- Agents remember corrections and still violate them — 2026-06-12
- Lean Retrieval Beat Full Context by Ten Points — 2026-06-11
- Curate, Don't Accumulate: Lean Memory, Mergeable Code, Shared Context — 2026-06-10
- The week the benchmarks broke — 2026-06-09
- Agents Get Graded on Process, Not Just Pass/Fail — 2026-06-09
- Weekly: the orchestration stack consolidates — 2026-06-08
Open models
Open-weight model releases and the ecosystem around running, fine-tuning, and evaluating them.
tagged in 17 of 82 digest issues, most recently 2026-07-13
Reliability
Failure modes, recovery, and durable state for agent systems running unattended or at production scale.
tagged in 4 of 82 digest issues, most recently 2026-07-11 · 10 papers · 1 explorer section
Papers
- ReAct: Synergizing Reasoning and Acting in Language Models — Interleaves chain-of-thought reasoning with tool actions so a model can plan, query external sources, and self-correct — reducing hallucination on decision tasks.
- Is Multi-Agent Debate (MAD) the Silver Bullet? Empirical Analysis in Code Summarization & Translation — Structured multi-agent debate yields minimal-to-inconsistent gains over a strong single-agent baseline on software-engineering tasks.
- Why Do Multi-Agent LLM Systems Fail? (MAST failure taxonomy) — 14 failure modes in 3 categories — specification issues, inter-agent misalignment, task verification — built with an LLM-as-judge pipeline at Cohen's κ=0.88.
- λ_A: A Typed Lambda Calculus for LLM Agent Composition — Well-formedness / termination guarantees for agent composition via a typed calculus.
- TraceFix: Repairing Agent Coordination Protocols with TLA+ Counterexamples — Uses TLA+ counterexamples to repair coordination protocols.
- SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work? — A 5-bucket failure taxonomy over 526 agent-attributable failures: implementation failure (41.6%) and timeout (31.4%) dominate, then reward hacking (15.4%), premature termination (7.6%), and poor self-verification (4.0%); long context degrades behavior actively — pass rate falls monotonically with consecutive-duplicate run length (claude-code 41.9%→3.2%) and compaction/summarizer trials pass at 0% vs 8.9% without.
- XFlow: An Executable Protocol Programming System for Reliable Multi-Agent Workflows — The paper makes the failure taxonomy explicit before scaling fan-out: a single agent's hallucinated/malformed/misinterpreted output becomes shared state and corrupts downstream decisions. Concrete failure modes observed: tau3-bench baseline reaches the right end state via an invalid path (treats a parameter clarification as authorization to mutate); CorpusQA fails on the interpretation rule not the retrieved value; SWE-bench baseline submits a patch after local validation already failed. The cloud-edge fan-out is gated: edge workers see only assigned chunks, write only declared outputs, and must pass schema + coverage checks before entering global state.
- AgentArmor: A Framework, Evaluation, & Mitigation of Coding Agent Failures — Decomposes non-adversarial coding-agent failure into three sequential points — forming the correct target (underspecification), pursuing it (capability error), and executing it through the harness (stochastic sampling, context decay) — with a chain-rule risk P(unsafe)=1-(1-f1)(1-f2)(1-f3) and scenarios that isolate each stage, cross-cut by four active modes (greenfield, editing, deployment, monitoring) over 8 scenarios, 20 environments, and 59 transcript templates at n>=500 across three frontier models.
- RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents — Names a failure taxonomy for the lab-to-production gap — fragile fixes (patches pass tests but leave latent bugs), token waste (trial-and-error instead of a planned approach), false confidence (never abstaining on impossible/ambiguous tasks), and broken intermediates (codebase left broken between steps). Baseline ReAct agents fail badly against it: no baseline abstained on any of the 6 impossible tasks, and even disciplined agents abstained correctly only 62% of the time. Recovery is the hardest mode and the one scaffolding does NOT fix — smallest gain of all five pillars, token-waste cut only 34%, doom loops persist when root cause isn't in the error message; the authors conclude recovery may need architectural changes beyond configuration-level frameworks.
- Beyond the Strongest LLM: Multi-Turn Multi-Agent Orchestration vs Single LLMs
Retrieval-augmented generation
Grounding model output in retrieved evidence: dense retrieval foundations, agentic search loops, and RAG system design.
tagged in 0 of 82 digest issues
Scientific search
Discovery over scholarly literature: citation graphs, semantic search, and agentic research assistants.
tagged in 0 of 82 digest issues
Synthetic data
Model-generated training and evaluation data: generation pipelines, quality control, and contamination risk.
8 papers · 1 explorer section · tagged in 0 of 82 digest issues
Papers
- Two Tales of Persona in LLMs: A Survey of Role-Playing and Personalization — Surveys the persona concept across role-playing and personalization, distinguishing how personas are constructed, conditioned, and evaluated in LLMs.
- Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models — Proposes evaluating synthetic-data generators by the Quality–Diversity–Complexity (QDC) makeup of their output, finding quality drives in-distribution generalization, diversity drives OOD generalization, and complexity helps both — with explicit quality–diversity trade-offs.
- Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces — Introduces OmniBehavior, the first user-simulation benchmark built entirely from real-world traces, and uses it to show LLM simulators converge to a 'positive average person' (hyper-activity, persona homogenization, Utopian bias), losing individual differences and long-tail behaviors.
- UniToolCall: Unifying Tool-Use Representation, Data, and Evaluation for LLM Agents — A unified tool-learning pipeline that builds a 22k+ tool pool and 390k+ instances by combining 10 public datasets with structurally controlled synthetic trajectories (single/multi-hop, single/multi-turn, serial/parallel), adds an Anchor Linkage mechanism for cross-turn dependencies, and a QAOA evaluation representation.
- Graph2Counsel: Clinically Grounded Synthetic Counseling Dialogue Generation from Client Psychological Graphs — Generates clinically grounded synthetic multi-turn counseling dialogues from structured client psychological graphs, encoding the underlying clinical reasoning rather than just masking real utterances, to sidestep confidentiality constraints on real data.
- EngramaBench: Evaluating Long-Term Conversational Memory with Structured Graph Retrieval — A synthetic long-term memory benchmark of 5 personas, 100 multi-session conversations, and 150 queries spanning factual recall, cross-space integration, temporal reasoning, adversarial abstention, and emergent synthesis, holding the answering model fixed (GPT-4o) to isolate memory architecture.
- A Survey on LLM-based Conversational User Simulation — A dedicated survey organizing the design space of LLM-based conversational user simulators (persona conditioning, goal/intent modeling, behavioral realism, evaluation of simulators themselves).
- VeriSim: A Configurable Framework for Evaluating Medical AI Under Realistic Patient Noise — A configurable simulation framework that injects realistic patient 'noise' (incomplete, inconsistent, distracting user behavior) into evaluation, exposing that strong static-benchmark scores collapse under realistic interaction.
Test-time compute
Spending inference-time computation to improve results: reasoning-intensive retrieval, reranking, and search at query time.
5 papers · 1 explorer section · tagged in 0 of 82 digest issues
Verification
Checking that an agent's output actually satisfies the task: test oracles, property checks, and validation of generated artifacts.
tagged in 2 of 82 digest issues, most recently 2026-06-27