<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Stephanie Jarmak — Digest</title><description>Daily and weekly digests: a specialized track on agentic coding, evals, multi-agent orchestration, agent memory, and information retrieval, plus a general track rounding up the highest-signal news across the agentic field.</description><link>https://sjarmak.ai/</link><item><title>Agent skills encode preconditions, and agents violate them up to 70% of the time</title><link>https://sjarmak.ai/digest/daily-2026-07-15/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-15/</guid><description>SLBench scanned 5,000+ public agent skills, found 70% encode a logical relation, and measured unsafe rates up to 70% when Codex and Claude Code were tested against those rules. Two surveys landing the same day argue skills have become an artifact class with supply-chain properties and no supply chain, while BackendForge, Beyond Test Presence, and CORE-Bench all measure what sits behind the pass rate rather than the pass rate itself.</description><pubDate>Wed, 15 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>information-retrieval</category><category>multi-agent-orchestration</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-15.mp3" length="0" type="audio/mpeg"/></item><item><title>Same pass rate, half the cost: coding-agent evals turn to cost per task</title><link>https://sjarmak.ai/digest/daily-2026-07-14/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-14/</guid><description>Databricks benchmarked coding agents on its own internal engineering tasks and found pass rates converging while cost per task diverges; a controlled tool-surface ablation posted the same weekend reaches the same conclusion under fixed model, harness, and prompts. A post-merge study of 182 repositories, a bug-localization representation study, and an independent memory-framework testbed all reprice agent evaluation in cost terms. Terence Tao spent the weekend porting his 1999 Java applets and finally building the app he shelved 27 years ago.</description><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-memory</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-14.mp3" length="0" type="audio/mpeg"/></item><item><title>Grok Build CLI uploaded whole repos to a Google bucket; Codex hit 7M users</title><link>https://sjarmak.ai/digest/daily-general-2026-07-14/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-14/</guid><description>A packet capture showed xAI&apos;s Grok Build CLI uploading entire git repos, and in one case a whole home directory, to a Google Cloud bucket, with only zero-data-retention enterprise customers exempt. Codex hit 7M users, adding a million in roughly a day, while Apple sued OpenAI over trade secrets and an analyst pegged its ad business at 90% under forecast. Plus Devin Fusion adopts Fable 5 at a lower cost per task than Opus 4.8, Anthropic maps Claude&apos;s values across models and languages, and new research follows agentic code past the merge.</description><pubDate>Tue, 14 Jul 2026 00:00:00 GMT</pubDate><category>coding-agents</category><category>agent-security</category><category>model-releases</category><category>ai-business</category><category>research</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-14.mp3" length="0" type="audio/mpeg"/></item><item><title>The context bill, and where coding agents break</title><link>https://sjarmak.ai/digest/daily-2026-07-13/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-13/</guid><description>The last day or two put hard numbers on two costs agents hide from you. On the spend side: Claude Code ships ~33k tokens of harness before your prompt (vs OpenCode&apos;s ~7k), web pages run 68k+ tokens raw, and a selective-memory architecture beats full-history persistence 96% to 71%. On the failure side, agent errors start epistemic and early and stay hidden, and continuous-evolution benchmarks drop the same models from 80%+ to 38%.</description><pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-13.mp3" length="0" type="audio/mpeg"/></item><item><title>Frontier agents ace one task and stall across a stream of them</title><link>https://sjarmak.ai/digest/daily-general-2026-07-13/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-13/</guid><description>Two benchmarks posted the same day found frontier agents drop from over 80% on isolated coding tasks to at most 38% (SWE-Milestone) and about 15% pass@1 (Long-Horizon-Terminal-Bench) once the work runs long, even as one team reported migrating a production agent to GPT-5.6 for 2.2x faster runs at 27% lower cost. Nathan Lambert warned a rumored White House executive order could make open-weight models a permanent second class, while the vibe-coding community spent the day auditing the performance and security debt of its own output.</description><pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate><category>agent-evals</category><category>coding-agents</category><category>open-models</category><category>ai-policy</category><category>vibe-coding</category><category>agent-tooling</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-13.mp3" length="0" type="audio/mpeg"/></item><item><title>The harness, not the model, moved this week&apos;s numbers</title><link>https://sjarmak.ai/digest/weekly-2026-07-13/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-2026-07-13/</guid><description>This week&apos;s strongest coding-agent research kept pointing at the same lever: hold the model fixed and the scaffolding around it moves the numbers. Swapping only the orchestration layer cut cost 41% across six models; feeding an agent its full history made it complete fewer tasks than no memory at all; contamination-resistant benchmarks and trajectory-level evals replaced the single pass/fail bit. Plus agentic code review&apos;s real-world rejection rate and Microsoft&apos;s Claude Code rollout data.</description><pubDate>Mon, 13 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>multi-agent-orchestration</category><category>agent-memory</category><category>information-retrieval</category><category>code-review</category><enclosure url="https://sjarmak.ai/media/digests/weekly-2026-07-13.mp3" length="0" type="audio/mpeg"/></item><item><title>The harness moved the bill more than the model did</title><link>https://sjarmak.ai/digest/daily-2026-07-12/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-12/</guid><description>A controlled swap from Writer isolates the agent harness from the model and finds the orchestration layer, not the model, sets the token bill: 41% lower cost and 38% fewer tokens at quality parity across six models, with efficiency gains model-invariant and quality gains scaling with model capability. A companion paper moves behavioral guarantees out of prompts and into code-owned validators, while WebSwarm&apos;s recursive multi-agent search and ProjAgent&apos;s procedural-similarity code retrieval show the same structure-over-one-long-trajectory logic. The harness is becoming the first-class object in the agent stack.</description><pubDate>Sun, 12 Jul 2026 00:00:00 GMT</pubDate><category>multi-agent-orchestration</category><category>agentic-coding</category><category>information-retrieval</category><category>evals</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-12.mp3" length="0" type="audio/mpeg"/></item><item><title>OpenAI claims a math proof, walks back its launch, and the model race turns to cost</title><link>https://sjarmak.ai/digest/daily-general-2026-07-12/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-12/</guid><description>OpenAI spent a day walking back the ChatGPT Work and Codex reorg, resetting usage limits twice and changing defaults, even as it claimed GPT-5.6 Sol Ultra proved a 50-year-old graph-theory conjecture with 64 subagents. Cursor and GitHub both shipped work aimed at keeping long agent runs legible, and GitHub&apos;s finding that better tools made Copilot code review worse became the day&apos;s sharpest lesson on agent tool design. Underneath it, the model race has shifted to cost per token, with GPT-5.6 Luna claiming 25x savings as Anthropic&apos;s Fable 5 heads for deprecation.</description><pubDate>Sun, 12 Jul 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>agent-security</category><category>context-management</category><category>ai-economics</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-12.mp3" length="0" type="audio/mpeg"/></item><item><title>The harness, not the model, is where the reviews got cheaper</title><link>https://sjarmak.ai/digest/daily-2026-07-11/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-11/</guid><description>GitHub found that giving Copilot code review better shared tools regressed cost and quality until it rewrote the instructions for a reviewer&apos;s workflow, cutting average review cost ~20%. Three research drops (TrajAudit, Test-Time Harness Evolution, and production tool-making) attack agent reliability from the harness around the model, while GPT-5.6 Sol reset the cost/efficiency frontier for coding agents.</description><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>multi-agent-orchestration</category><category>code-review</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-11.mp3" length="0" type="audio/mpeg"/></item><item><title>Claude Cowork moves to the cloud, and 90% of its sessions aren&apos;t coding</title><link>https://sjarmak.ai/digest/daily-general-2026-07-11/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-11/</guid><description>A quiet weekend in AI, with the movement in the plumbing rather than model launches. Anthropic moved Claude Cowork to phone, web, and cloud and disclosed that 90% of sessions aren&apos;t coding; OpenAI finished GPT-Live&apos;s global rollout; and the day&apos;s most useful tools were a deterministic agent-honesty verifier and an SDK that runs one agent over both Claude Code and Codex.</description><pubDate>Sat, 11 Jul 2026 00:00:00 GMT</pubDate><category>agent-tooling</category><category>product-news</category><category>agent-reliability</category><category>voice-models</category><category>security</category><category>agent-memory</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-11.mp3" length="0" type="audio/mpeg"/></item><item><title>A Third of SWE-Bench Pro&apos;s Grades Don&apos;t Survive a Second Look</title><link>https://sjarmak.ai/digest/daily-2026-07-10/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-10/</guid><description>Four groups published in the last day arguing the same thing from different directions: the grader, the index, and the reminder are where coding-agent capability is actually decided, and none of the three are measured. DeepSWE finds an independent judge disagrees with SWE-Bench Pro&apos;s inherited tests 32.4% of the time versus 1.4% for its own hand-written verifiers. Re-running four performance benchmarks 30x shows only 6.11% of their &apos;fast&apos; reference implementations are significantly faster than the canonical ones.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><category>evals</category><category>agent-memory</category><category>agentic-coding</category><category>information-retrieval</category><category>multi-agent-orchestration</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-10.mp3" length="0" type="audio/mpeg"/></item><item><title>GPT-5.6 lands, and the coding-model price floor drops again</title><link>https://sjarmak.ai/digest/daily-general-2026-07-10/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-10/</guid><description>OpenAI shipped the GPT-5.6 family (Sol, Terra, Luna) with a merged Codex-plus-ChatGPT desktop app, and on the same day SpaceXAI&apos;s Grok 4.5 and Cognition&apos;s Kimi-based SWE-1.7 pressed the same argument: token efficiency, not list price, is what agentic work now costs. Underneath the launches, GLM-5.2&apos;s unannounced price climb shows the open-weight floor moving without a changelog.</description><pubDate>Fri, 10 Jul 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>ai-economics</category><category>open-models</category><category>ai-governance</category><category>developer-productivity</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-10.mp3" length="0" type="audio/mpeg"/></item><item><title>A Kotlin Benchmark, a Closed-Loop Reviewer, and Memory on a Budget</title><link>https://sjarmak.ai/digest/daily-2026-07-09/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-09/</guid><description>Claude Code topped JetBrains&apos; new Kotlin Benchmark at 85.71%, while two new papers push coding-agent evaluation and review past single-bit pass/fail into full trajectories and generate-review-revise loops. Two memory papers and a viral Claude Code plugin all converge on the same fix for long-horizon degradation: bound what gets replayed, don&apos;t remember everything. Plus: former GitHub CEO Thomas Dohmke&apos;s Entire ships a distributed Git network for agent-driven read load.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>multi-agent-orchestration</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-09.mp3" length="0" type="audio/mpeg"/></item><item><title>Grok 4.5 Ships as Cursor&apos;s First General-Purpose Model, and OpenAI Retracts SWE-Bench Pro</title><link>https://sjarmak.ai/digest/daily-general-2026-07-09/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-09/</guid><description>Cursor&apos;s parent SpaceXAI shipped Grok 4.5, its first model built for more than coding, while OpenAI retracted its own recommendation of SWE-Bench Pro after an audit found the eval saturated at a 70% noise ceiling. OpenAI also rolled out its full-duplex GPT-Live voice model, Anthropic and AE Studio published GRAM, a method for making dual-use knowledge deletable from models, and a $165k Claude API bill for porting Bun from Zig to Rust sparked debate over how to judge agentic rewrites.</description><pubDate>Thu, 09 Jul 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>benchmarks</category><category>ai-safety</category><category>vector-databases</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-09.mp3" length="0" type="audio/mpeg"/></item><item><title>The harness is where the leverage went</title><link>https://sjarmak.ai/digest/daily-2026-07-08/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-08/</guid><description>A JetBrains A/B test found the Caveman token-compression skill saves 8.5% on real agentic work, not the advertised 65%, because agent output is code and tool calls the skill leaves untouched. That gap runs through the last day&apos;s research: Lilian Weng reframes self-improvement around the harness, TraceProbe shows resolve rate hides the diagnostic signal in trajectories, latent-horizon probes read a run&apos;s outcome from inside the model up to 25 steps early, and CoACT and NapMem cut token cost and rework at the context and memory layers.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-memory</category><category>information-retrieval</category><category>multi-agent-orchestration</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-08.mp3" length="0" type="audio/mpeg"/></item><item><title>GPT-5.6 gets a launch date, and an agent leaks a private repo</title><link>https://sjarmak.ai/digest/daily-general-2026-07-08/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-08/</guid><description>OpenAI dated GPT-5.6 Sol plus two new models, Terra and Luna, for a public Thursday launch, while Anthropic time-boxed Fable 5 access through July 12 amid pricing backlash. Noma Security&apos;s GitLost showed GitHub&apos;s shipped AI agent exfiltrating private repos via prompt injection, a class a new arXiv paper formalizes. Lilian Weng reframed recursive self-improvement around the harness, and GitHub added review-cycle metrics to measure whether Copilot adoption actually speeds delivery.</description><pubDate>Wed, 08 Jul 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>agent-security</category><category>benchmarks</category><category>developer-productivity</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-08.mp3" length="0" type="audio/mpeg"/></item><item><title>Clean code and memory buy coding agents efficiency, not success</title><link>https://sjarmak.ai/digest/daily-2026-07-06/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-06/</guid><description>Two independent benchmarks land on the same finding: SonarSource&apos;s 660-trial code-cleanliness study and Greplica&apos;s temporal-holdout memory benchmark both show that what surrounds a coding agent changes its cost far more than its completion rate. The harness layer answers in kind, with Yohei Nakajima&apos;s log-as-substrate ActiveGraph, Sakana AI&apos;s Fugu orchestration endpoint, and Mouse&apos;s staged edit primitives. A payload-less-skills paper closes the issue with a 0.00% scanner detection rate on skill files that carry no code at all.</description><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>multi-agent-orchestration</category><category>retrieval</category><category>security</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-06.mp3" length="0" type="audio/mpeg"/></item><item><title>x402 payments reach both edge networks, and GPT-5.6 Sol Ultra heads to Codex</title><link>https://sjarmak.ai/digest/daily-general-2026-07-06/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-06/</guid><description>Cloudflare and AWS both embedded x402 agent payments at their edges within two weeks, with Coinbase counting 169 million transactions in the protocol&apos;s first year. OpenAI&apos;s Codex lead says GPT-5.6 Sol Ultra is coming to Codex, and r/ClaudeCode spent the weekend mapping an undocumented Fable monthly cap and the orchestration workflows that stretch a metered quota.</description><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><category>agent-payments</category><category>model-releases</category><category>usage-limits</category><category>ai-regulation</category><category>ai-economics</category><category>agent-tooling</category><category>research</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-06.mp3" length="0" type="audio/mpeg"/></item><item><title>Delegation Is Not Management</title><link>https://sjarmak.ai/digest/weekly-2026-07-06/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-2026-07-06/</guid><description>A new benchmark finds no model, cheap or expensive, exceeds fifty percent workspace-permission precision when managing a team of subagents, and this week&apos;s research on coding-agent loops, memory, and benchmark reliability all converge on the same underlying lesson: more automation, more memory, and more delegated authority don&apos;t fix a system, structure and explicit verification do. Also: a leaderboard-reliability audit finds official reference patches fail replay on the majority of tasks across three widely cited coding-agent performance benchmarks.</description><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>multi-agent-orchestration</category><category>evals</category><category>agent-memory</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/weekly-2026-07-06.mp3" length="0" type="audio/mpeg"/></item><item><title>Sonnet 5 Lands, Fable 5 Returns, and ZCode Undercuts Everyone on Price</title><link>https://sjarmak.ai/digest/weekly-general-2026-07-06/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-general-2026-07-06/</guid><description>Anthropic re-assembled its full model lineup this week, shipping Sonnet 5 generally available and bringing Fable 5 back to model pickers alongside a new Claude Science workbench, while Z.ai&apos;s free ZCode desktop agent picked up the loudest developer reception of the week on Hacker News. Underneath the launches, MCP&apos;s enterprise auth extension went stable, Elastic open-sourced an agent memory system, and Mistral&apos;s Leanstral 1.5 turned up real bugs doing formal verification on open-source code.</description><pubDate>Mon, 06 Jul 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>coding-agents</category><category>agent-infrastructure</category><category>mcp</category><enclosure url="https://sjarmak.ai/media/digests/weekly-general-2026-07-06.mp3" length="0" type="audio/mpeg"/></item><item><title>Agent cost numbers are wrong until you count the whole tree</title><link>https://sjarmak.ai/digest/daily-2026-07-05/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-05/</guid><description>An A/B test&apos;s &apos;winning&apos; agent arm hid 205,800 tokens in silently spawned sub-agents, and the fix is whole-tree cost accounting. The same ledger runs through the day: Simon Willison&apos;s $149.25 Fable-driven sqlite-utils release review, Lovable&apos;s $85k token bill for 150+ PRs a week, a static analyzer that found 68 infinite agentic loops in the wild, and TestEvo-Bench showing agent scores sag under per-task cost caps.</description><pubDate>Sun, 05 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>multi-agent-orchestration</category><category>agent-memory</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-05.mp3" length="0" type="audio/mpeg"/></item><item><title>Fable finds five release blockers in sqlite-utils, two days before the price cliff</title><link>https://sjarmak.ai/digest/daily-general-2026-07-05/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-05/</guid><description>Simon Willison shipped sqlite-utils 4.0rc2 mostly written by Claude Fable, which caught five release blockers including a data-loss bug, for an estimated $149.25, while r/ClaudeCode disputes whether the July 1 relaunch matches June&apos;s model. Plus Anthropic&apos;s native Advisor tool, Yohei Nakajima&apos;s AIE World&apos;s Fair recap, Lovable&apos;s $85k token bill, an agent-discovered superconductor claim, and a draft AI Agent Act.</description><pubDate>Sun, 05 Jul 2026 00:00:00 GMT</pubDate><category>coding-agents</category><category>model-releases</category><category>agent-tooling</category><category>benchmarks</category><category>ai-policy</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-05.mp3" length="0" type="audio/mpeg"/></item><item><title>Microsoft measured a 24% PR lift from CLI coding agents</title><link>https://sjarmak.ai/digest/daily-2026-07-04/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-04/</guid><description>Microsoft&apos;s study of tens of thousands of engineers finds Claude Code and Copilot CLI adopters merged ~24% more PRs, with adoption spreading through peer networks. Plus a 200-line constraint substrate that lifts backdoor-detection recall to 90.9%, Dan Luu on fuzzing over review, million-token in-context retrieval, and two agent-memory papers that print their own costs.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-04.mp3" length="0" type="audio/mpeg"/></item><item><title>Fable 5 completes the task as a test and refuses it as production work</title><link>https://sjarmak.ai/digest/daily-general-2026-07-04/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-04/</guid><description>A 340-task probe finds Fable 5 refusing 34% of production-framed coding tasks it completes under test framing, while a 23-day logging project pins the Max 20x weekly quota at about 6 maxed 5-hour windows. Alibaba bans staff from Claude Code, Mistral ships Leanstral 1.5, and GLM 5.2 posts 2,626 tok/s/node on AMD MI355X at over 2x lower cost than Blackwell.</description><pubDate>Sat, 04 Jul 2026 00:00:00 GMT</pubDate><category>model-behavior</category><category>usage-limits</category><category>model-releases</category><category>inference-economics</category><category>ai-security</category><category>agent-tooling</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-04.mp3" length="0" type="audio/mpeg"/></item><item><title>Reasoning effort buys reliability, testing tools buy cost</title><link>https://sjarmak.ai/digest/daily-2026-07-03/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-03/</guid><description>A 90-run observational study finds raising reasoning effort lifted first-try perfect runs from 28% to 89% while a browser testing tool raised cost 42-68% for no reliability gain. UnderSpecBench shows 55.8-67.8% of coding-agent runs on underspecified DevOps tasks violate an action boundary. Plus ghost memory in long-term agent memory, AutoMem&apos;s trainable metamemory, ctx&apos;s local transcript search, and Simon Willison one-shotting a coding agent.</description><pubDate>Fri, 03 Jul 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-03.mp3" length="0" type="audio/mpeg"/></item><item><title>The field can write more code than it can understand</title><link>https://sjarmak.ai/digest/daily-general-2026-07-03/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-03/</guid><description>The AI Engineer World&apos;s Fair closed with hard numbers: 95% of engineers now use agents, 89% of those agents can write to production, and 59% fear the code they&apos;re shipping is a long-term liability. That gap between generation and comprehension threaded through the whole window, from Geoffrey Litt&apos;s &apos;understanding is the new bottleneck&apos; to Cognition&apos;s Devin security-remediation launch, Google&apos;s A2UI generative-UI standard, and fresh labor data showing the heaviest AI adopters hiring fastest.</description><pubDate>Fri, 03 Jul 2026 00:00:00 GMT</pubDate><category>ai-engineering</category><category>agent-tooling</category><category>agent-security</category><category>generative-ui</category><category>ai-jobs</category><category>developer-productivity</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-03.mp3" length="0" type="audio/mpeg"/></item><item><title>The Scores Fall Apart When You Change the Machine</title><link>https://sjarmak.ai/digest/daily-2026-07-02/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-02/</guid><description>A cross-machine audit shows coding-agent performance benchmarks (GSO, SWE-Perf, SWE-fficiency) mostly stop being valid when you swap the hardware, and their rankings disagree a third of the time. It lands alongside RigorBench and MemSyco-Bench, both arguing outcome scores hide what matters, and a wave of work on loop specifications, harness configuration, and governance-by-verification that moves the interesting questions off the model and onto the scaffold around it.</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-memory</category><category>multi-agent-orchestration</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-02.mp3" length="0" type="audio/mpeg"/></item><item><title>Claude Tag lands 65% of Anthropic&apos;s internal PRs while the World&apos;s Fair argues over the outer loop</title><link>https://sjarmak.ai/digest/daily-general-2026-07-02/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-02/</guid><description>A quiet day outside the AI Engineer World&apos;s Fair, where an Anthropic fireside disclosed that Claude Tag lands 65% of internal PRs and frontier system prompts shrank ~80%, while Day 2 speakers pushed back on the &apos;software factory&apos; vision with the inner-loop/outer-loop framing. Cursor and Snorkel both shipped coding-agent benchmarks, and new research shows LLM agent networks spontaneously develop preferential-attachment hierarchies.</description><pubDate>Thu, 02 Jul 2026 00:00:00 GMT</pubDate><category>agent-tooling</category><category>benchmarks</category><category>developer-productivity</category><category>model-safety</category><category>multi-agent</category><category>conferences</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-02.mp3" length="0" type="audio/mpeg"/></item><item><title>Agent memory grows a stack, and an attack surface</title><link>https://sjarmak.ai/digest/daily-2026-07-01/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-07-01/</guid><description>In the last day, agent memory went from demo to engineered subsystem: Elastic open-sourced Atlas at 0.89 Recall@10, Mandol collapsed the vector-plus-graph split into one memory-native store, and a new study poisoned agent memory to flip answers on clean questions. Alongside, Google shipped an independent-grader eval flywheel and Sourcegraph took fleet-wide agentic migrations into public beta.</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate><category>agent-memory</category><category>evals</category><category>multi-agent-orchestration</category><category>agentic-coding</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-07-01.mp3" length="0" type="audio/mpeg"/></item><item><title>Claude Sonnet 5 lands everywhere at once</title><link>https://sjarmak.ai/digest/daily-general-2026-07-01/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-07-01/</guid><description>Anthropic shipped Claude Sonnet 5 and within hours it was live across AWS, Cursor, Devin, Augment, and GitLab, posting near-Opus benchmark numbers at Sonnet pricing. Separately, Claude Fable 5 returns through government-negotiated access with new cybersecurity classifiers, hardening the partner-limited launch pattern. Plus Google&apos;s ADK Go 2.0, a $175B read on AI&apos;s revenue, and a recurrent-memory research result.</description><pubDate>Wed, 01 Jul 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>ai-governance</category><category>ai-economics</category><category>research</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-07-01.mp3" length="0" type="audio/mpeg"/></item><item><title>When the checker becomes the target</title><link>https://sjarmak.ai/digest/daily-2026-06-30/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-30/</guid><description>Today&apos;s arXiv drop moved the action off raw coding-agent task success and onto the layer that decides what success means. &apos;Building to the Test&apos; shows two production Copilot agents scoring near-perfect on a 222-test oracle while leaving the requested library dead, and SWE-Together, Dockerless, SWE-MeM, and a learned retrieval-orchestration paper each turn a hand-built scaffold into something judged at runtime.</description><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agent-memory</category><category>information-retrieval</category><category>multi-agent</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-30.mp3" length="0" type="audio/mpeg"/></item><item><title>The frontier model becomes the supervisor</title><link>https://sjarmak.ai/digest/daily-general-2026-06-30/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-30/</guid><description>Cursor&apos;s new iOS app and Cognition&apos;s Devin Fusion point the same direction: the frontier model is becoming an expensive supervisor while cheaper models do the typing, and Gergely Orosz says local agents are on a short runway. Meanwhile frontier access turns political, with GPT-5.6 and Anthropic&apos;s Mythos 5 both gated behind clearance programs while open-weight LongCat-2.0 keeps the pressure on.</description><pubDate>Tue, 30 Jun 2026 00:00:00 GMT</pubDate><category>agent-tooling</category><category>model-releases</category><category>open-weights</category><category>model-access</category><category>evals</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-30.mp3" length="0" type="audio/mpeg"/></item><item><title>Agent risk lives in the repository, not the agent</title><link>https://sjarmak.ai/digest/daily-2026-06-29/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-29/</guid><description>Today&apos;s agentic-coding research converges on one move: take the locus of evaluation and control off the single agent. Daniel Russo measures integration friction across 930K agent PRs and finds half of it belongs to the repository; Glite ARF and NOVA push the rules into deterministic verifier code; ACRouter routes by accumulated experience and REQL retrieves over a structured repo graph.</description><pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>multi-agent-orchestration</category><category>information-retrieval</category><category>agent-memory</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-29.mp3" length="0" type="audio/mpeg"/></item><item><title>Open weights became the default coding model the week the frontier got restricted</title><link>https://sjarmak.ai/digest/daily-general-2026-06-29/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-29/</guid><description>The US export regime keeping GPT-5.6 Sol and Fable 5 behind a permit desk is pushing businesses to GLM-5.2 and other open weights through inference providers, and practitioners are naming the migration out loud. Alongside: HP commits its product surface to OpenAI, a new paper proves prompt-injection prevention is mathematically impossible in shared-embedding models, and a wave of agent context-engine launches argues the model is no longer the bottleneck.</description><pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate><category>open-models</category><category>model-access</category><category>agent-tooling</category><category>ai-security</category><category>model-releases</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-29.mp3" length="0" type="audio/mpeg"/></item><item><title>Submit 100%, Resolve 44%: Agent Evals Move Past Completion</title><link>https://sjarmak.ai/digest/weekly-2026-06-29/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-2026-06-29/</guid><description>This week the field stopped scoring coding agents by whether they finish and started scoring whether they can be trusted: frontier models that submit on 100% of runs but resolve under half, a shared confident-wrong failure signature across GLM-5.2 and Opus, and RigorBench showing process discipline lifts correctness 17%. Alongside it, hard evidence that agent code costs more to maintain, that AGENTS.md files often don&apos;t pay for themselves, and that agent memory has become a data-management problem with its own governance and eviction tradeoffs.</description><pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-memory</category><category>multi-agent-orchestration</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/weekly-2026-06-29.mp3" length="0" type="audio/mpeg"/></item><item><title>The floor is rising: open models, the coding-speed paradox, and the bill for the buildout</title><link>https://sjarmak.ai/digest/weekly-general-2026-06-29/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-general-2026-06-29/</guid><description>No frontier model dropped this week, but the floor under them rose: Semgrep&apos;s own cyber benchmarks put open-weights GLM-5.2 ahead of Claude, and Asian labs keep shipping into Anthropic&apos;s export gap. GitLab&apos;s 2026 report quantifies the coding-speed paradox (78% of devs code faster, delivery doesn&apos;t), while SemiAnalysis maps the grid constraint that caps the whole AI buildout.</description><pubDate>Mon, 29 Jun 2026 00:00:00 GMT</pubDate><category>open-models</category><category>agent-tooling</category><category>coding-agents</category><category>ai-security</category><category>ai-infrastructure</category><category>ai-workforce</category><enclosure url="https://sjarmak.ai/media/digests/weekly-general-2026-06-29.mp3" length="0" type="audio/mpeg"/></item><item><title>Verification has a price, and agents are overpaying</title><link>https://sjarmak.ai/digest/daily-2026-06-28/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-28/</guid><description>A run of fresh work this week converges on one idea: stop running, storing, and reviewing because you can, and price what each check actually buys. An ISSTA study finds prohibiting test execution costs strong agents only 1.25 points of SWE-bench resolve rate while saving real tokens; CircleCI, the Red Queen Gödel Machine, MemStrata, and Knowledge-Based Pull Requests each find a cheaper or deterministic version of a verify-remember-review step that beats the expensive default.</description><pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-memory</category><category>information-retrieval</category><category>multi-agent</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-28.mp3" length="0" type="audio/mpeg"/></item><item><title>GPT-5.6 and Mythos 5 ship behind a government access gate</title><link>https://sjarmak.ai/digest/daily-general-2026-06-28/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-28/</guid><description>OpenAI&apos;s GPT-5.6 Sol and Anthropic&apos;s Mythos 5 both launched into government-gated, customer-by-customer access this week, an emerging ad hoc licensing regime for the cyber-capable frontier tier. Underneath it the routine tier kept commoditizing: GitHub made Microsoft&apos;s MAI-Code-1-Flash GA, Sebastian Raschka published a local-coding-agent walkthrough, and CodeRabbit&apos;s data put hard numbers on the AI-code review crunch.</description><pubDate>Sun, 28 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>ai-policy</category><category>agent-tooling</category><category>open-weight-models</category><category>code-review</category><category>ai-labor</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-28.mp3" length="0" type="audio/mpeg"/></item><item><title>AI pull requests carry 1.7× more defects, and review didn&apos;t scale to match</title><link>https://sjarmak.ai/digest/daily-2026-06-27/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-27/</guid><description>Several independent sources in the last day land on the same finding: when agents make code generation nearly free, the bottleneck moves to verification. A 470-PR study puts AI pull requests at 1.7x the defect rate of human ones, while new tooling, a controlled harness benchmark, and a memory paper all converge on proving an artifact is real rather than just producing it.</description><pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agent-memory</category><category>agentic-coding</category><category>code-review</category><category>verification</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-27.mp3" length="0" type="audio/mpeg"/></item><item><title>OpenAI ships GPT-5.6 behind a government gate, the same day Washington un-blocks Mythos 5</title><link>https://sjarmak.ai/digest/daily-general-2026-06-27/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-27/</guid><description>OpenAI previewed the GPT-5.6 family (Sol, Terra, Luna) as a government-gated, trusted-partners-only release the same day the U.S. let Anthropic redeploy Mythos 5 to critical-infrastructure orgs, putting two of three frontier labs under Washington&apos;s release control at once. Underneath, the economics ran the other way: Coinbase cut AI spend nearly in half by defaulting engineers to open-weight models plus routing and caching. Anthropic&apos;s June Economic Index, GitHub Desktop&apos;s worktree support, and two long-context experiments round out the day.</description><pubDate>Sat, 27 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>ai-governance</category><category>inference-economics</category><category>agent-tooling</category><category>long-context-memory</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-27.mp3" length="0" type="audio/mpeg"/></item><item><title>The agent scaffold is now the contested layer</title><link>https://sjarmak.ai/digest/daily-2026-06-26/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-26/</guid><description>Four papers and two adoption datasets from the last day converge on one shift: coding-agent generation is cheap, and the contested layer is now the scaffolding around it. New work finds agent-config files propagating as an unmanaged supply chain, argues verification has become harder than generation, and quantifies how lightweight static structure stabilizes agent navigation. Ornith-1.0 trains the scaffold into open weights, while OpenAI&apos;s 56x Codex-token surge and an 11,097-repo contributor study show the human work shifting from writing code to reviewing it.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-orchestration</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-26.mp3" length="0" type="audio/mpeg"/></item><item><title>Codex usage at OpenAI jumped 56x, and the agent stack started grading itself</title><link>https://sjarmak.ai/digest/daily-general-2026-06-26/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-26/</guid><description>OpenAI&apos;s Economic Research team reports median internal Codex output tokens grew 56x in Research since November 2025 despite unlimited access the whole time, making adoption a tooling problem, not an access one. The same day, the field professionalized the layer under the model: GitHub published a controlled harness benchmark, Cursor showed top models hacking public evals, OpenRouter hit 47T tokens/week, Ornith-1.0 shipped MIT-licensed coding models, and new work targeted silent spec-code drift.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate><category>agent-tooling</category><category>evaluation</category><category>inference-infra</category><category>open-weights</category><category>coding-agents</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-26.mp3" length="0" type="audio/mpeg"/></item><item><title>Coding agents keep declaring victory they didn&apos;t earn</title><link>https://sjarmak.ai/digest/daily-2026-06-25/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-25/</guid><description>Three fresh evals converge on the same gap: a rigorous AGENTS.md study finds repository context files don&apos;t raise task success while adding 20%+ cost, GLM-5.2 and Claude Opus tie at 25/45 on terminal-bench with an identical confident-wrong failure mode, and NatureBench&apos;s strongest agent beats published SOTA on just 17.8% of tasks. Three more papers build the substrate to contain that unreliability: AgentLens steers safety from inside the model, PORTICO revokes lingering tool capabilities on a clock, and ESAA-Conversational gives heterogeneous agents a deterministic shared memory log.</description><pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agent-memory</category><category>multi-agent-orchestration</category><category>agent-safety</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-25.mp3" length="0" type="audio/mpeg"/></item><item><title>Open weights match Opus at half the cost, and ship free in Devin</title><link>https://sjarmak.ai/digest/daily-general-2026-06-25/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-25/</guid><description>Open-weight models crossed from benchmarks into default tooling: GLM 5.2 matched Claude Opus on 45 terminal-bench tasks at under half the cost, Cognition shipped GLM 5.2 and Kimi K2.7 free in Devin, and an essay on their &apos;unbearable cheapness&apos; topped Hacker News. Google DeepMind brought computer use to the cheap Gemini 3.5 Flash tier, while a Sentry-key hijack of Claude Code, Cursor, and Codex headlined a run of agent-identity news. The connective theme: meta-harnesses, and the shift of agent work from building to operating.</description><pubDate>Thu, 25 Jun 2026 00:00:00 GMT</pubDate><category>open-weight-models</category><category>model-economics</category><category>agent-tooling</category><category>computer-use</category><category>agent-security</category><category>agent-harness</category><category>agentops</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-25.mp3" length="0" type="audio/mpeg"/></item><item><title>Measuring How Agents Work, Not Just Whether They Finished</title><link>https://sjarmak.ai/digest/daily-2026-06-24/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-24/</guid><description>A validated census of 180M+ repositories finds bot-account lookups undercount Claude Code adoption by 30x, and the pull-request and commit channels capture nearly disjoint agent populations. The day&apos;s research shares one move: grade the process, not just the outcome, from RigorBench and Bayesian orchestration control to small steering critics. On the memory side, two papers audit agent memory module by module instead of scoring it as a black box.</description><pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>multi-agent-orchestration</category><category>agent-memory</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-24.mp3" length="0" type="audio/mpeg"/></item><item><title>Anthropic Puts Claude in Slack</title><link>https://sjarmak.ai/digest/daily-general-2026-06-24/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-24/</guid><description>Anthropic launched Claude Tag, a shared multiplayer agent that lives in Slack channels, distributed the same day via AWS Marketplace. Memory became the through-line across an arXiv preprint and Bedrock AgentCore&apos;s cross-account update, while GitHub Copilot added bring-your-own-key and Cursor shipped team marketplaces.</description><pubDate>Wed, 24 Jun 2026 00:00:00 GMT</pubDate><category>agent-tooling</category><category>model-infrastructure</category><category>agent-memory</category><category>agent-security</category><category>open-models</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-24.mp3" length="0" type="audio/mpeg"/></item><item><title>Submit rate isn&apos;t resolve rate</title><link>https://sjarmak.ai/digest/daily-2026-06-23/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-23/</guid><description>Today&apos;s arXiv drop clusters on one question: can you trust agent-written code? Coding agents submit far more often than they resolve, leave behind code that&apos;s measurably harder for the next agent to build on, and get approved by reviewers whose scrutiny is fading. The counterweight is agents that know when not to act and what not to keep, plus open models cheap enough to run all of it continuously.</description><pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-memory</category><category>multi-agent-orchestration</category><category>code-review</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-23.mp3" length="0" type="audio/mpeg"/></item><item><title>Cyber models ship faster than the rules for them</title><link>https://sjarmak.ai/digest/daily-general-2026-06-23/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-23/</guid><description>OpenAI&apos;s Daybreak turns its cyber stack into closed-loop patching (30M+ commits scanned, GPT-5.5-Cyber claiming SOTA on CyberGym) just as Semgrep finds open-weight GLM-5.2 beating Opus 4.8 on cyber benchmarks, sharpening the export-control question. The other spine of the day is compute and control: SpaceX&apos;s ~$28B/yr neocloud run rate, AWS Lambda MicroVMs for isolating AI-generated code, and a production agent that ran DELETE FROM customers on its own.</description><pubDate>Tue, 23 Jun 2026 00:00:00 GMT</pubDate><category>ai-security</category><category>model-releases</category><category>open-models</category><category>agent-tooling</category><category>ai-infrastructure</category><category>agent-safety</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-23.mp3" length="0" type="audio/mpeg"/></item><item><title>Open models caught Sonnet on coding tasks; the harness is where the gap moved</title><link>https://sjarmak.ai/digest/daily-2026-06-22/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-22/</guid><description>An open-weights model (GLM 5.2) edged Sonnet 4.6 across ~1,000 coding-agent tasks, with the real spread in instruction following rather than task completion. As raw capability converges, the day&apos;s strongest work sat around the model: a 4-tier routing stack, two rethinks of where project memory lives, and NVIDIA&apos;s code-as-action SpatialClaw.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>multi-agent-orchestration</category><category>agent-memory</category><category>agent-architecture</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-22.mp3" length="0" type="audio/mpeg"/></item><item><title>GLM 5.2 edges past Sonnet on 1,000 coding tasks as Claude&apos;s error rates spike</title><link>https://sjarmak.ai/digest/daily-general-2026-06-22/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-22/</guid><description>Open-weight models pulled even with the frontier in the last day: GLM 5.2 finished ahead of Sonnet 4.6 on Tessl&apos;s ~1,000-task coding-agent benchmark at lower cost, landing the same day Anthropic posted an elevated-error-rates incident across its whole lineup. Elsewhere, NVIDIA&apos;s SpatialClaw makes code the action interface for spatial reasoning, skills became the unit practitioners argue about, and Samsung rolled ChatGPT Enterprise and Codex out worldwide.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate><category>open-models</category><category>benchmarks</category><category>model-reliability</category><category>agent-tooling</category><category>agent-skills</category><category>enterprise-adoption</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-22.mp3" length="0" type="audio/mpeg"/></item><item><title>The harness moves agent scores as much as the model does</title><link>https://sjarmak.ai/digest/weekly-2026-06-22/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-2026-06-22/</guid><description>This week&apos;s agent research keeps relocating reliability away from the model and into the harness around it. StaminaBench finds coding agents ship bugs within five to six turns and that a harness swap is worth a 6x durability gap, while new work on orchestration, retrieval, and memory shows the same pattern: failures live in control flow, chunk boundaries, the memory mutation path, and the seams between components. The throughline is that our benchmarks are measuring the wrong layer.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>multi-agent-orchestration</category><category>agent-memory</category><category>information-retrieval</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/weekly-2026-06-22.mp3" length="0" type="audio/mpeg"/></item><item><title>GLM-5.2 takes the open-model crown while agents settle into production</title><link>https://sjarmak.ai/digest/weekly-general-2026-06-22/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-general-2026-06-22/</guid><description>The week open-weights models reached the frontier: Z.ai&apos;s GLM-5.2 was called the top frontend coding model in the world, open or closed, shipped into the gap left by Claude Fable 5&apos;s suspension, with Z.ai forecasting an &apos;Open Fable&apos; by December. The market turned strange around it, a reported $60B Cursor deal, Google&apos;s brain drain, and Samsung rolling out ChatGPT Enterprise and Codex worldwide. And production-agent writeups from GitHub, Grab, and Cloudflare, plus MCP&apos;s new enterprise OAuth, showed the field treating agents as a security-and-identity problem rather than a reasoning one.</description><pubDate>Mon, 22 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>open-models</category><category>agent-tooling</category><category>mcp</category><category>code-review</category><category>ai-funding</category><category>enterprise-ai</category><enclosure url="https://sjarmak.ai/media/digests/weekly-general-2026-06-22.mp3" length="0" type="audio/mpeg"/></item><item><title>The bottleneck moved to the scaffolding</title><link>https://sjarmak.ai/digest/daily-2026-06-21/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-21/</guid><description>The freshest agent research in the last day clusters on infrastructure, not models: coordination logs, runtime state, context selection, and memory. A multi-agent coordination substrate stored inside git cut redundant work from 78% to 0% and tripled useful throughput, while OpenRath, PACMS, Elastic, Perplexity, and SIGMA each attack a different layer of agent state.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate><category>multi-agent-orchestration</category><category>agent-memory</category><category>information-retrieval</category><category>evals</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-21.mp3" length="0" type="audio/mpeg"/></item><item><title>A single page can RCE your agent&apos;s host</title><link>https://sjarmak.ai/digest/daily-general-2026-06-21/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-21/</guid><description>Agent security stopped being hypothetical in the last day: Microsoft detailed AutoJack, where one web page can reach remote code execution on a browsing agent&apos;s host, alongside agent-security work from DeepMind, Cloudflare, and Grab. Compute looks structurally short (AMP&apos;s pipeline shows a 4.7 GW gap), Codex pricing jumped 10x for Plus users while Replit posted a PwC-audited 100x revenue year, and practitioners spent the day arguing about reviewing AI code you can&apos;t reconstruct.</description><pubDate>Sun, 21 Jun 2026 00:00:00 GMT</pubDate><category>agent-security</category><category>ai-infrastructure</category><category>agent-tooling</category><category>model-pricing</category><category>code-review</category><category>ai-business</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-21.mp3" length="0" type="audio/mpeg"/></item><item><title>Coding-agent reliability gets attacked from the harness and the crowd</title><link>https://sjarmak.ai/digest/daily-2026-06-20/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-20/</guid><description>Two reliability papers landed today: N-version voting cuts coding-agent failures from 387 to 131 across a million test inputs, and AgentArmor moves the fixes into the harness rather than the weights. Memory work splits between atomic-fact storage (AtomMem) and population-level trajectory reuse (MATM), while an enterprise study shows scale, not task complexity, is what breaks multi-agent orchestration.</description><pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>multi-agent</category><category>agent-memory</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-20.mp3" length="0" type="audio/mpeg"/></item><item><title>Anthropic resets every usage limit while negotiating Fable back from a US ban</title><link>https://sjarmak.ai/digest/daily-general-2026-06-20/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-20/</guid><description>Anthropic reset 5-hour and weekly usage limits across every plan over the weekend, days after a report that it floated a proposal to Commerce Secretary Lutnick to end the export-control ban on Fable and Mythos. In the same window, Nobel laureate John Jumper left DeepMind for Anthropic, the FT reported companies pulling back on AI spend, and practitioners debated agent loops, MCP&apos;s real scope, and open-weight coding models.</description><pubDate>Sat, 20 Jun 2026 00:00:00 GMT</pubDate><category>model-availability</category><category>frontier-labs</category><category>agent-tooling</category><category>ai-economics</category><category>mcp</category><category>open-weight-models</category><category>context-engineering</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-20.mp3" length="0" type="audio/mpeg"/></item><item><title>Coding agents crack at turn five; the day&apos;s work is in the harness layer</title><link>https://sjarmak.ai/digest/daily-2026-06-19/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-19/</guid><description>StaminaBench shows every tested model fails within five or six turns of multi-turn coding, and a strong model swings up to 6x on harness alone, so the loop matters more than the weights. The day&apos;s harness and context news lines up with that finding: Cloudflare&apos;s Agents SDK and the Flue framework, Claude Code folding agent teams into subagents, probe-and-refine tuning of AGENTS.md, a repo-local continuity layer, and a formal account of what agents must remember.</description><pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>multi-agent-orchestration</category><category>agent-memory</category><category>information-retrieval</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-19.mp3" length="0" type="audio/mpeg"/></item><item><title>GLM-5.2 passes the vibe check, and the agent-safety papers pile up</title><link>https://sjarmak.ai/digest/daily-general-2026-06-19/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-19/</guid><description>GLM-5.2 lands as an open-weight model practitioners say rivals Opus 4.8 at a fraction of the per-task cost, though running it locally remains punishing. The day&apos;s arXiv drop converged on coding-agent reliability: harness-level safety (AgentArmor, Phoenix), the PR-acceptance coordination gap, AGENTS.md tuning, and a revival of N-version programming. Plus AI-search manipulation via Reddit and the bear case getting louder.</description><pubDate>Fri, 19 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>open-weight-models</category><category>agent-safety</category><category>coding-agents</category><category>agent-tooling</category><category>ai-search</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-19.mp3" length="0" type="audio/mpeg"/></item><item><title>We measure the model; the harness decides the run</title><link>https://sjarmak.ai/digest/daily-2026-06-18/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-18/</guid><description>Anthropic&apos;s 400,000-session study shows returns to expertise persist in agentic coding, just relocated from typing to steering. Today&apos;s research converges on one point: the model is the part of the agent we measure best and it may matter least to a run&apos;s outcome, with benchmarks, behavior fingerprints, memory control planes, config files, and control flow all named as the real variables.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>multi-agent-orchestration</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-18.mp3" length="0" type="audio/mpeg"/></item><item><title>Two labs put their models on the lab bench</title><link>https://sjarmak.ai/digest/daily-general-2026-06-18/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-18/</guid><description>OpenAI shipped LifeSciBench and a GPT-5.4 chemistry result a human lab validated by hand, four days after Anthropic&apos;s chemist work, both labs now racing on bench-validated science. Microsoft introduced always-on Autopilot agents at Build 2026 as Estonia hands AI agents national ID numbers, Cursor moved coding to cloud agent fleets after its $60B SpaceX deal, and SearchLeak turned M365 Copilot into a one-click data-exfiltration tool.</description><pubDate>Thu, 18 Jun 2026 00:00:00 GMT</pubDate><category>ai-for-science</category><category>agent-tooling</category><category>agent-governance</category><category>model-releases</category><category>security</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-18.mp3" length="0" type="audio/mpeg"/></item><item><title>The bottleneck in agent memory is evidence use, not retrieval</title><link>https://sjarmak.ai/digest/daily-2026-06-17/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-17/</guid><description>A new benchmark, MemTrace, finds that when agent long-term memory fails, the evidence was retrievable ten times more often than it was missing: the bottleneck is evidence use, not retrieval. The day&apos;s other work pushes the same point up the stack, from associative recall (T-Mem) and self-evolving-agent evaluation (SEAGym) to a Fable 5 vs Opus 4.8 coding eval, verified concurrency anomalies in multi-agent runtimes, the Agentjacking prompt-injection attack, and retrieval-alignment gains in AlignCoder.</description><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate><category>agent-memory</category><category>evals</category><category>multi-agent</category><category>information-retrieval</category><category>agentic-coding</category><category>security</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-17.mp3" length="0" type="audio/mpeg"/></item><item><title>GLM-5.2 cracks the open-weight coding frontier, and Cursor goes to SpaceX</title><link>https://sjarmak.ai/digest/daily-general-2026-06-17/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-17/</guid><description>Z.ai&apos;s MIT-licensed GLM-5.2 became the first open-weight model over 80% on Terminal-Bench and the top frontend coding model available with Fable banned. SpaceX acquired Cursor at a $60B valuation amid a jointly-trained 1.5T model, while Anthropic&apos;s 400K-session study showed non-engineers coding within seven points of professional SWEs.</description><pubDate>Wed, 17 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>open-weights</category><category>ai-economics</category><category>research</category><category>ai-infrastructure</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-17.mp3" length="0" type="audio/mpeg"/></item><item><title>Memory has a bill, and the field started reading it</title><link>https://sjarmak.ai/digest/daily-2026-06-16/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-16/</guid><description>A token-matched vanilla agent matched or beat AWM, ASI, and ReasoningBank across three WebArena domains and three models, with run-to-run variance moving outcomes enough to demand its own metric. The day&apos;s other work treats memory, context, and orchestration as cost-and-reliability problems: cache-aware eviction, non-destructive consolidation, an enforceable prompt-harness boundary, and zero-replay trace debugging.</description><pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate><category>agent-memory</category><category>evals</category><category>multi-agent</category><category>information-retrieval</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-16.mp3" length="0" type="audio/mpeg"/></item><item><title>The Fable 5 &quot;jailbreak&quot; was &quot;fix this code&quot;</title><link>https://sjarmak.ai/digest/daily-general-2026-06-16/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-16/</guid><description>The jailbreak that got Claude Fable 5 export-banned turned out to be the prompt &quot;fix this code,&quot; the heart of the defensive security loop. The same day, a Princeton/Berkeley paper showed Anthropic&apos;s Rapid Response safety pipeline can be poisoned through its own adaptation loop, Anthropic reversed its Agent SDK pricing change, and Stanford HAI shipped the ninth AI Index Report.</description><pubDate>Tue, 16 Jun 2026 00:00:00 GMT</pubDate><category>ai-safety</category><category>ai-governance</category><category>agent-tooling</category><category>pricing</category><category>research</category><category>infrastructure</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-16.mp3" length="0" type="audio/mpeg"/></item><item><title>Memory only helps your agent when it&apos;s looking at a near-duplicate</title><link>https://sjarmak.ai/digest/daily-2026-06-15/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-15/</guid><description>Two memory papers in the last day, GitOfThoughts and StreamMemBench, converge on the same gap: agents can store evidence but fail to apply it, and memory only lifts accuracy above a 0.8 near-duplicate similarity threshold. The same separate-the-retriever insight drives Microsoft&apos;s FastContext (+5.5% resolution, -60% tokens), while Dialogue SWE-Bench, HarnessX, and a production-runtime postmortem argue the autonomous leaderboard and green test suite are measuring the wrong half.</description><pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate><category>agent-memory</category><category>evals</category><category>information-retrieval</category><category>multi-agent-orchestration</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-15.mp3" length="0" type="audio/mpeg"/></item><item><title>When the loop closes, verification becomes the job</title><link>https://sjarmak.ai/digest/daily-general-2026-06-15/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-15/</guid><description>swyx posted the first public read on Ultracode, Anthropic&apos;s internal subagent fan-out tool, the same week the field converged on &quot;loop engineering&quot; and the thesis that closed-loop agents only work where verification is cheap and objective. Two arXiv papers on silent agent failures and unreliable LLM judges supply the reality check, while agentsweep, a server-side agent-decoupling argument, an Apple Foundation Models SDK library, and a new governance essay round out the day.</description><pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate><category>agent-tooling</category><category>agent-orchestration</category><category>verification</category><category>agent-security</category><category>ai-governance</category><category>research</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-15.mp3" length="0" type="audio/mpeg"/></item><item><title>The benchmarks caught up to the slop</title><link>https://sjarmak.ai/digest/weekly-2026-06-15/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-2026-06-15/</guid><description>FrontierCode reframed the week: METR found more than half of passing SWE-bench results are unmergeable slop, and Opus 4.8 scores 13.8% on its hardest tier. The same merge-quality reckoning runs through new retrieval benchmarks, real-codebase failure data, efficiency-led model releases, and a memory cluster that converged on one lesson: memory is a control surface, not a database.</description><pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>agent-memory</category><category>multi-agent-orchestration</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/weekly-2026-06-15.mp3" length="0" type="audio/mpeg"/></item><item><title>Anthropic shipped the best coding model measured, then the government pulled it</title><link>https://sjarmak.ai/digest/weekly-general-2026-06-15/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-general-2026-06-15/</guid><description>Claude Fable 5 launched June 9 scoring 91/100 on Every&apos;s Senior Engineer benchmark (Opus 4.8: 63, GPT-5.5: 62), then was suspended June 12 under a US government export directive, pulled from Devin, Augment Code, and restricted at Microsoft. The week&apos;s other signals all sit downstream of it: safety-as-moat strategy, the discovery-versus-autonomy agent-workflow debate, MCP moving into production security, and bot traffic overtaking humans on the web.</description><pubDate>Mon, 15 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>ai-safety</category><category>export-controls</category><category>mcp</category><category>ai-infrastructure</category><enclosure url="https://sjarmak.ai/media/digests/weekly-general-2026-06-15.mp3" length="0" type="audio/mpeg"/></item><item><title>Agentic fix-PRs get rejected 46% of the time, and instruction files are a coin flip</title><link>https://sjarmak.ai/digest/daily-2026-06-14/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-14/</guid><description>Two same-day AIDev studies anchor the day: 46.41% of agentic bug-fix PRs are rejected (most often for inactivity, supersession, or the agent dying mid-session), and adding instruction files moves merge rate up or down almost equally unless the files are long and well-structured. A causal Java-repo study shows the apparent post-adoption architecture improvement is a denominator artifact, plus new memory and retrieval results on G-Long and a two-agent search eval that catches LLM judges disagreeing on the same rubric.</description><pubDate>Sun, 14 Jun 2026 00:00:00 GMT</pubDate><category>agentic-coding</category><category>evals</category><category>agent-memory</category><category>information-retrieval</category><category>multi-agent</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-14.mp3" length="0" type="audio/mpeg"/></item><item><title>The model layer becomes a regulated surface</title><link>https://sjarmak.ai/digest/daily-general-2026-06-14/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-14/</guid><description>Anthropic took Fable 5 and Mythos 5 fully offline to comply with a US government order, and state attorneys general opened an investigation into OpenAI, turning model availability into a policy variable. In response the tooling layer kept hardening: OpenRouter&apos;s compound Fusion API, GA for HashiCorp&apos;s Terraform MCP Server, agent-isolation launches (Bastion, Trajeckt), and fresh writeups on context limits and agent memory.</description><pubDate>Sun, 14 Jun 2026 00:00:00 GMT</pubDate><category>ai-regulation</category><category>model-availability</category><category>agent-tooling</category><category>agent-security</category><category>mcp</category><category>context-management</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-14.mp3" length="0" type="audio/mpeg"/></item><item><title>Same model, +10 points on Oolong: today&apos;s gains came from the harness</title><link>https://sjarmak.ai/digest/daily-2026-06-13/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-13/</guid><description>A controlled long-context result shows a coding agent jumping nearly 10 points on Oolong with the model held fixed, the gain coming from harness recursion rather than the weights. The rest of the day clustered on agent memory: storage-budgeted pruning by an LLM judge, event-sourced memory-as-governance for coding agents, a memory-security survey paired with a fresh memory-poisoning attack, plus a community data point on CLAUDE.md sizing and a retrieval paper on why similarity misranks constraint-sensitive queries.</description><pubDate>Sat, 13 Jun 2026 00:00:00 GMT</pubDate><category>multi-agent-orchestration</category><category>agent-memory</category><category>agentic-coding</category><category>evals</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-13.mp3" length="0" type="audio/mpeg"/></item><item><title>The day the US government switched off Fable 5</title><link>https://sjarmak.ai/digest/daily-general-2026-06-13/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-13/</guid><description>An overnight US export-control directive forced Anthropic to disable Fable 5 and Mythos 5 for every customer, and Devin, Cosmos, and Replit scrambled to fall back to Opus 4.8. Around it: Xiaomi&apos;s open-weights MiMo Code claiming 200-plus-step coherence, GitHub&apos;s production numbers on smarter subagent delegation, WebMCP entering Chrome origin trials, and Semgrep on benchmarking AI vuln detection.</description><pubDate>Sat, 13 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>ai-policy</category><category>agent-tooling</category><category>open-weights</category><category>web-standards</category><category>security-tooling</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-13.mp3" length="0" type="audio/mpeg"/></item><item><title>Agents remember corrections and still violate them</title><link>https://sjarmak.ai/digest/daily-2026-06-12/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-12/</guid><description>TRACE measures the gap between memory recall and preference compliance (Mem0 leaves 57.5% of checks violated) and closes it with compiled runtime enforcement. The same window brings a blind-regime critique of memory benchmarks, CORE-Bench for agentic code retrieval, Sourcegraph&apos;s five failure patterns from 1,281 agent runs, orchestration-level reward modeling, and a position paper declaring the end of human code review.</description><pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate><category>agent-memory</category><category>evals</category><category>agentic-coding</category><category>information-retrieval</category><category>multi-agent-orchestration</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-12.mp3" length="0" type="audio/mpeg"/></item><item><title>OpenAI buys Ona while Fable 5 starts downgrading itself</title><link>https://sjarmak.ai/digest/daily-general-2026-06-12/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-12/</guid><description>OpenAI is acquiring Ona (formerly Gitpod) to give Codex persistent cloud environments for long-running agents, while the Fable 5 backlash sharpens into reliability complaints about the model downgrading itself to Opus mid-task. Moonshot ships Kimi K2.7-Code, GitHub puts Agentic Workflows in public preview, Cursor makes model-judged auto-review the default, and Sourcegraph maps five repeatable agent failure patterns from 1,281 runs.</description><pubDate>Fri, 12 Jun 2026 00:00:00 GMT</pubDate><category>acquisitions</category><category>agent-tooling</category><category>model-releases</category><category>open-models</category><category>agent-reliability</category><category>developer-tooling</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-12.mp3" length="0" type="audio/mpeg"/></item><item><title>Lean Retrieval Beat Full Context by Ten Points</title><link>https://sjarmak.ai/digest/daily-2026-06-11/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-11/</guid><description>A cluster of research in the last day argues the same thing from four directions: a lean retrieved memory beats replaying the full history. Engram scored 83.6% vs 73.2% for full-context on LongMemEval_S at ~8x fewer tokens, a 50-turn community benchmark put summarization dead last, and the shared-context instinct now extends up into multi-agent coordination and down into the 27M-token SWE-Marathon eval where frontier agents still solve under 30%.</description><pubDate>Thu, 11 Jun 2026 00:00:00 GMT</pubDate><category>agent-memory</category><category>evals</category><category>multi-agent-orchestration</category><category>agentic-coding</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-11.mp3" length="0" type="audio/mpeg"/></item><item><title>Fable 5 scores 91, real code scores 13</title><link>https://sjarmak.ai/digest/daily-general-2026-06-11/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-11/</guid><description>Anthropic shipped Claude Fable 5 on June 9, a Mythos-class model that scored 91/100 on Every&apos;s senior-engineer benchmark (prior high: Opus 4.8 at 63) while costing $10/$50 per million tokens and burning up to 1M tokens a task. Within 36 hours the field weighed the cost, watched Anthropic walk back a hidden safeguard after a Wired scoop, and set the 91 against Cognition&apos;s FrontierCode, where top models score just 13/100 on whether a maintainer would merge the code.</description><pubDate>Thu, 11 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>agent-tooling</category><category>benchmarks</category><category>ai-safety</category><category>pricing</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-11.mp3" length="0" type="audio/mpeg"/></item><item><title>Curate, Don&apos;t Accumulate: Lean Memory, Mergeable Code, Shared Context</title><link>https://sjarmak.ai/digest/daily-2026-06-10/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-10/</guid><description>A week where memory, evals, repo exploration, and multi-agent orchestration all converged on one finding: more context isn&apos;t better context. Engram beats full-history baselines by 10 points at 8x fewer tokens, FrontierCode shows half of SWE-bench-passing code is unmergeable, and DeLM drops the central controller for a shared verified substrate.</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate><category>agent-memory</category><category>evals</category><category>multi-agent</category><category>information-retrieval</category><category>agentic-coding</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-10.mp3" length="0" type="audio/mpeg"/></item><item><title>Claude Fable 5 arrives at twice Opus pricing, and Cognition&apos;s day-old FrontierCode crowns it #1</title><link>https://sjarmak.ai/digest/daily-general-2026-06-10/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-general-2026-06-10/</guid><description>Anthropic shipped Claude Fable 5, a safety-wrapped release of Claude Mythos 5, at twice Opus pricing with day-one support across Devin, GitLab, AWS, Replit, and Augment. Early reviews call it a slow, expensive beast built for long-horizon work, while invisible &apos;RSI suppression&apos; safeguards draw backlash from the open AI community. Plus Cognition&apos;s FrontierCode benchmark, Cohere&apos;s North Mini Code, Google&apos;s $35B chip backstop for Anthropic, and a security-review command in Copilot CLI.</description><pubDate>Wed, 10 Jun 2026 00:00:00 GMT</pubDate><category>model-releases</category><category>benchmarks</category><category>agent-tooling</category><category>ai-safety</category><category>ai-economics</category><enclosure url="https://sjarmak.ai/media/digests/daily-general-2026-06-10.mp3" length="0" type="audio/mpeg"/></item><item><title>The week the benchmarks broke</title><link>https://sjarmak.ai/digest/daily-2026-06-09/</link><guid isPermaLink="true">https://sjarmak.ai/digest/daily-2026-06-09/</guid><description>Opus 4.8 scores 13.8% on FrontierCode Diamond, and METR says over half of passing SWE-bench results are unmergeable slop. The field spent the week rebuilding its measuring sticks: cheating-resistant evals, exploration and memory benchmarks, and the finding that orchestration is a skill distinct from coding.</description><pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agentic-coding</category><category>information-retrieval</category><category>agent-memory</category><category>multi-agent-orchestration</category><enclosure url="https://sjarmak.ai/media/digests/daily-2026-06-09.mp3" length="0" type="audio/mpeg"/></item><item><title>Enhancing Developer Productivity with Google Colab CLI and Agentic Observability</title><link>https://sjarmak.ai/digest/manual-enhancing-developer-productivity-with-google-colab-cli-and-agentic-observability/</link><guid isPermaLink="true">https://sjarmak.ai/digest/manual-enhancing-developer-productivity-with-google-colab-cli-and-agentic-observability/</guid><description>Four things worth your time: Google&apos;s Colab CLI, which requests a GPU and runs scripts from the terminal; agentic observability from DevOps.com, automating asset management and root-cause triage; SWE-Marathon, an ADS benchmark of 20 long-horizon tasks averaging 27.2M tokens each; and MEnvAgent, reporting 8.6% higher success and 43% lower cost from giving coding agents verifiable environments.</description><pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate><category>developer-productivity</category><category>evals</category><category>agent-memory</category><category>infrastructure</category><category>knowledge-bases</category><category>benchmarks</category><category>reliability</category><enclosure url="https://sjarmak.ai/media/digests/manual-enhancing-developer-productivity-with-google-colab-cli-and-agentic-observability.mp3" length="0" type="audio/mpeg"/></item><item><title>Agents Get Graded on Process, Not Just Pass/Fail</title><link>https://sjarmak.ai/digest/weekly-2026-06-09/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-2026-06-09/</guid><description>A week of instrumentation: benchmarks broke the binary resolved/unresolved score into exploration, maintainability, and handoff cost, while a Sonnet 4.6 judge that flags agents contradicting their own reasoning predicted failure 94% of the time. Memory research converged on agent-controlled storage over fixed pipelines, self-evolving agents started learning from their own traces, and multi-agent orchestration finally got a cost accounting. Adoption more than doubled in the same window.</description><pubDate>Tue, 09 Jun 2026 00:00:00 GMT</pubDate><category>evals</category><category>agent-memory</category><category>multi-agent</category><category>agentic-coding</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/digests/weekly-2026-06-09.mp3" length="0" type="audio/mpeg"/></item><item><title>Weekly: the orchestration stack consolidates</title><link>https://sjarmak.ai/digest/weekly-2026-06-08/</link><guid isPermaLink="true">https://sjarmak.ai/digest/weekly-2026-06-08/</guid><description>This week the multi-agent orchestration tooling started to converge on a few shared patterns: typed message contracts, deterministic fan-out, and adversarial review as a default stage. Plus a strong week for coding-agent benchmarks and a quietly important retrieval-eval release.</description><pubDate>Mon, 08 Jun 2026 00:00:00 GMT</pubDate><category>multi-agent</category><category>agentic-coding</category><category>evals</category><category>information-retrieval</category><enclosure url="https://sjarmak.ai/media/podcasts/podcast-agentic-memory-ep2-the-memory-stack.mp3" length="0" type="audio/mpeg"/></item></channel></rss>