AI Radar Research

Daily research digest for developers — Wednesday, September 09 2026

arXiv

Memory as Infrastructure: Reliability Engineering for Persistent Agent Memory in Months-Long LLM-Assisted Development

This paper discusses the challenges and solutions for managing long-term memory in LLM coding agents that operate over extended periods and large codebases.

Why it matters: Understanding how to maintain reliable memory over long projects is crucial for developing robust AI coding tools.
arXiv

Grounded Skill Synthesis from Code at Scale for Agentic Intelligence

The paper presents a method for synthesizing reusable skills from code, enabling agents to acquire procedural knowledge that can be transferred across tasks.

Why it matters: Skill synthesis is essential for creating versatile AI coding agents that can adapt to new challenges.
arXiv

Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models

This study compares iterative edit-based generation and direct generation methods for LLMs in code editing, focusing on Flutter/Dart models.

Why it matters: Choosing the right generation method can significantly impact the efficiency and accuracy of AI-assisted code editing.
arXiv

Broken on Arrival: Silently Defective LLM Artifacts in Public Model Registries and How to Catch Them

This paper addresses the issue of defective LLM artifacts in public registries and proposes methods for detecting and preventing such defects.

Why it matters: Ensuring the integrity of AI models is vital for developers relying on public model registries.
arXiv

When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents

The paper evaluates the impact of long-term memory on the performance of tool-using LLM agents, using a new benchmark called MERIT.

Why it matters: Understanding when and how memory aids AI agents can optimize their design and functionality.
arXiv

AutoFyn Technical Report: Non-Parametric Expert Iteration for Long-Horizon Agents

AutoFyn introduces a non-parametric expert iteration approach for long-horizon agents, using persistent state updates rather than model weight changes.

Why it matters: This approach could lead to more efficient training of autonomous coding agents over long tasks.
arXiv

CriticGen: Generation-Aware Evaluation as Actionable Feedback

CriticGen proposes a fine-grained evaluation method for LLMs that provides actionable feedback, improving model generation quality.

Why it matters: Actionable feedback is crucial for refining AI coding tools and enhancing their output quality.
arXiv

Correct Tests Are Not Enough: Measuring and Training Oracle Conversion in Specification-Based Test Generation

This research highlights the importance of oracle conversion in test generation, emphasizing that correct tests alone are insufficient for robust software validation.

Why it matters: Improving test generation processes can lead to more reliable AI-assisted software development.
arXiv

Look Before You Prompt, and After: Scaffolding Human-AI Collaboration in Software Tutorial Creation

This paper explores how LLMs can assist in creating software tutorials, focusing on the collaborative process between humans and AI.

Why it matters: Enhancing human-AI collaboration can improve the quality and accessibility of educational resources in software engineering.
arXiv

When Agent Governance Helps

The paper discusses the design and evaluation of governed autotelic AI agent organizations, where agents pursue self-generated goals within set guardrails.

Why it matters: Understanding governance in AI agents is key to ensuring safe and aligned autonomous systems.
✉ Subscribe to daily research digest