AI Radar Research

Daily research digest for developers — Monday, July 13 2026

arXiv

GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning

This paper introduces GATS, a novel approach for improving the efficiency of agent planning by integrating graph-augmented tree search with layered world models.

Why it matters: GATS offers a more computationally efficient method for multi-step reasoning in AI coding agents, potentially reducing costs and improving performance.
arXiv

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading

This paper presents a new benchmark for evaluating AI agents on long-horizon tasks using dense reward-based grading to better assess their capabilities.

Why it matters: The benchmark provides a more comprehensive evaluation of AI coding systems, particularly in complex, long-duration tasks.
arXiv

ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning

ARCANA is a multi-agent framework designed to solve ARC AGI 2 tasks by decomposing them into perception, hypothesis generation, symbolic execution, and reflective refinement.

Why it matters: ARCANA demonstrates a structured approach to program synthesis, enhancing the reliability and efficiency of autonomous coding agents.
arXiv

SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation

SCATE proposes a method for supervising autonomous coding agents to generate tests more cost-effectively by addressing the issue of lazy generation.

Why it matters: SCATE improves the cost-effectiveness of test generation, making AI coding tools more practical for large-scale deployment.
arXiv

The Patchwork Problem in LLM-Generated Code

This paper discusses the 'patchwork problem' where LLM-generated code often appears correct but fails in deployment due to structural issues.

Why it matters: Understanding and addressing the patchwork problem is crucial for improving the reliability of AI-generated code.
arXiv

Programmable Agents Are Poor and Overconfident Judges of LLM-Generated Assertions

This study finds that programmers often misjudge the correctness of LLM-generated assertions, leading to potential overconfidence in AI-generated code.

Why it matters: The findings highlight the need for improved tools and training to better evaluate AI-generated code.
arXiv

Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation

This paper explores the use of small language models (SLMs) in place of larger models to reduce inference costs, using automated harness adaptation to maintain performance.

Why it matters: The approach offers a cost-effective alternative for deploying AI coding agents without sacrificing performance.
arXiv

AgentKGV: Agentic LLM-RAG Framework with Two-Stage Training for the Fact Verification of Knowledge Graphs

AgentKGV introduces a framework for fact verification in knowledge graphs using a two-stage training process to improve accuracy and reliability.

Why it matters: This framework enhances the reliability of AI systems in verifying large-scale data, crucial for maintaining data integrity.
arXiv

CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions

CogniConsole proposes a formal abstraction for managing inference-time control in LLM systems, aiming to improve reliability and interaction quality.

Why it matters: The approach enhances the reliability of LLM interactions, which is critical for developing dependable AI coding tools.
arXiv

HALO: Hybrid Adaptive Latent Reasoning for Language Models

HALO explores improving frozen pretrained language models with adaptive computation, enhancing their reasoning capabilities without extensive retraining.

Why it matters: This method provides a way to enhance LLM reasoning capabilities efficiently, benefiting AI coding tools that rely on complex reasoning.
✉ Subscribe to daily research digest