i
DATAIST
Back to feed

AI for coding

Models that write, read and fix code — from autocomplete to working through a task in a repository on their own.

60 articles

Anthropic prices Claude Haiku 5.5 to match GPT-6 Luna

Anthropic has released Claude Haiku 5.5 with API prices 90% below Haiku 4.5 for shorter requests, putting it at the same low-end rates as OpenAI’s GPT-6 Luna. The launch is aimed less at replacing larger models than at making a small one cheap enough to call often: Anthropic positions Haiku as a helper for Opus and Sonnet, while reserving harder coding work for those larger systems. That makes the pricing threshold—and what happens to real task costs beyond it—central to the pitch.

Microsoft’s Agent Lightning trains agents without rebuilding their pipelines

Microsoft Research Asia has released Agent Lightning v1.0, an open-source framework for training AI agents with reinforcement learning while keeping their existing pipelines intact. The framework is about 3,500 lines of code and puts a model proxy between the agent and the training system, rather than requiring developers to rebuild the agent inside a training framework. In a coding-agent example, training Qwen3.5-9B on about 6,000 samples raised its SWE-bench Verified Pass@1 score from 41.8% to 56.4%.

Meta’s Claude Code users fall as Anthropic moves into its market

Meta has cut the number of Claude Code users at the company from about 60,000 to 30,000, as it promotes its own coding tools and Anthropic expands into products that compete with its partners. The shift is notable not just for the scale of Meta’s spending on Claude Code—more than $105 million in a single 28-day period—but for the changing relationship between the companies: Anthropic is no longer just a supplier Meta needs to accommodate.

HackerRank makes AI literacy part of the coding interview

HackerRank is turning its coding interview into a work simulation: candidates solve a task in a real code repository with an AI assistant, while Chakra watches how they reason and asks follow-up questions. The shift matters because the company argues that AI has made a polished final answer less revealing. Its new product aims to assess judgment and a candidate’s ability to work with AI, not just whether they can reach the right result.

Google’s RRSI gives AI agents smaller gains on familiar tests

Google researchers have proposed a way to keep self-optimizing AI agents from learning the test set instead of getting better at the work. Their method, RRSI, changes how an agent’s software framework is revised and selected while leaving the underlying model untouched. In tests across coding, office work and engineering design, it delivered smaller gains on familiar tasks than competing optimizers, but improved performance on every unfamiliar benchmark the researchers tested.

Anthropic adds Mods to let developers change Claude Code from within

Anthropic has added Mods to Claude Code, a way to change how the coding tool behaves from inside the product. The extensions can run in the command-line and desktop apps, with partial support in the VS Code extension. Anthropic has also published example Mods on GitHub and introduced its first official plugin, which flags information in Claude’s answers that a user might have missed. The key trade-off is that Mods run with the user’s permissions and are not sandboxed.

MIT’s SIFT cuts the cost of evaluating coding agents

MIT’s SIFT makes the search for better coding agents cheaper by letting a language model compare candidate versions before they face a full benchmark. The method combines those judgments with quick checks and asynchronous testing, so teams can keep exploring while expensive evaluations run. Its results suggest that a model’s code-based judgment can sometimes pick a stronger agent than a small test set can—but the benchmark still has the final say.

OpenAI’s cheaper Sol still faces the refusal problem

OpenAI is trying to make its newest models refuse fewer harmless requests without loosening protections for more capable systems. GPT-6.1 Sol, introduced alongside that effort, is said to approach GPT-6 Astra on coding and computer use while costing one-fifth as much per token. But developers say safeguards still interrupt ordinary work, from robotics to cybersecurity. That tension matters more as cheaper models make repeated refusals a recurring cost in engineers’ time.

OpenAI releases GPT-6.1 Sol, a cheaper model nearing Astra

OpenAI has released GPT-6.1 Sol, a cheaper model that the company says comes close to GPT-6 Astra on several measures. It is aimed at difficult work such as coding, debugging, document understanding and multi-step workflows. The launch also narrows the gap between Sol and Astra on factual errors, while leaving the more consequential question unresolved: how much confidence should users place in those safety claims when OpenAI’s reported tests are the evidence?

OpenAI prices GPT-6.1 Sol at a fifth of delayed Astra

OpenAI is releasing GPT-6.1 Sol as a lower-cost alternative to its delayed flagship, GPT-6.1 Astra. The company says Sol comes close to Astra on AI-assisted coding, computer use and office work, while costing a fifth as much. That makes the launch more than a cheaper model release: with Astra held back over safety concerns, OpenAI is putting a less capable but more deployable option in front of customers.

OpenAI adds security reviews and faster decisions to Codex and API

OpenAI is extending Codex from a coding assistant into a broader working layer for developers: shared cloud environments, security reviews, and agents that can operate software through its API. Alongside those tools, the company is adding a fast-response API and a premium mode that can generate up to 300 tokens per second. The announcements matter less as a single product launch than as a map of where OpenAI wants its developer stack to go: from writing code toward coordinating work, checking it, and acting on it.

Jev targets AI decisions with a cheaper alternative to generation

AI systems often generate paragraphs when the application needs only a label, score or routing decision. TypeSafe’s Jev is built for that narrower job: it selects among structured outcomes in parallel rather than decoding a response token by token. The company says Jev is about 194 times faster and 445 times cheaper than generative models on its own workflow evaluations. Those figures make the pitch compelling, but they describe a particular kind of task—and TypeSafe says they may sit near the upper bound of real-world results.

Nvidia’s SoL-Pi cuts coding-agent tokens by optimizing the pipeline

Nvidia’s SoL-Pi cuts coding-agent token use by optimizing the control layer, not the model. The system searches for changes to the pipeline that connects an agent to its working environment, then tests them against tasks kept out of the search process. On EdgeBench, its most economical configuration used 49% fewer tokens than Pi while retaining 93.7% of Pi’s score. That makes the result less a claim about smarter models than a case for treating the agent’s operating logic as a major cost lever.

Supabase data exposure puts AI-built apps’ security defaults to the test

About 16,000 Supabase databases exposed some form of personal information, according to security firm UpGuard, which shared its findings with TechCrunch. The exposed records included names, addresses, phone numbers and passwords. The scale matters because Supabase has become a popular place to store data for apps built with AI coding tools—where a project can be easy to launch and still be left open to the internet.

David Heinemeier Hansson stops coding by hand

David Heinemeier Hansson, creator of Ruby on Rails, says he has stopped writing code manually. He sees the change as more than a productivity shift: AI agents are altering both the programmer’s role and the abstractions software is built around. Developers are becoming people who direct machine intelligence, he argues, even though the profession has not yet agreed on how that work should be done.

Fabrix.ai puts three Argos models behind Governed VibeOps

Fabrix.ai is positioning Governed VibeOps as a control layer for enterprise vibe coding, rather than a fix for vibe coding itself. Shailesh Manjrekar, the company’s director of AI marketing and strategy, says organizations need a way to control what coding agents can access, change and spend. Fabrix says it already has customers using the system in production since nearly the start of this year.

Lovable's annual revenue run rate tops $600 million

Lovable says its annual revenue run rate has passed $600 million as AI-assisted software development expands beyond code generation. The company is pitching itself less as a coding tool than as a place to build and operate products, and says two-thirds of Fortune 500 companies now use it. That shift matters because apps built on the platform collectively attract almost 1 billion views a month.

Xiaomi’s open MiMo models target the cost of AI agents

Xiaomi has released MiMo-V2.6-Pro, an MIT-licensed open model that the company says now leads open-weight systems on Artificial Analysis’ intelligence ranking. Alongside it comes MiMo-V2.6-Flash, a smaller model priced at roughly one-third of Pro’s API rates while staying close on several agent benchmarks. The release matters less as a single leaderboard result than as evidence of Xiaomi’s broader strategy: build an open stack for AI agents, from models and coding tools to training environments and reinforcement-learning infrastructure.

Grok 4.7 keeps token prices low while coding costs rise

SpaceXAI has released Grok 4.7 as a stronger coding and knowledge-work model while keeping its standard API price at $2/$6 per 1 million input and output tokens. The launch matters less for the headline rate than for the tradeoff underneath it: Grok 4.7 often scores better than Grok 4.6, but it may use far more tokens to finish a task. That can turn a cheaper model into a more expensive one once it reaches production workloads.

xAI launches Grok 4.7 at $2 per million input tokens

xAI has released Grok 4.7 as a coding and knowledge-work model, pricing it at $2 per million input tokens and $6 per million output tokens. The price is closer to Chinese models than to leading Western systems, but independent tests place Grok 4.7 well behind Claude Fable 5.1 and GPT-6. That makes the launch less a frontier-model challenge than a bet that lower operating cost can compensate for weaker performance.

How AI Agents Can Save Context and Avoid Failures

Autonomous coding agents can lose their way not because they cannot write code, but because they run out of room to remember what they are doing. The authors examine how the surrounding software harness—the tools and rules that guide an AI agent—affects its ability to solve long, complicated programming tasks. They compare ways to manage context, plan work, and choose actions, showing that the best setup depends on the model’s strengths and on how much memory it has available. The results reveal when planning prevents mistakes, when it mainly saves time and cost, and why simpler command-line control can sometimes work better than a larger toolbox. In this review we look at how these design choices shape an agent’s entire problem-solving path, and what they suggest for building coding systems that stay effective without wasting context.

G5 Labs wants natural-language intent to replace code as source

G5 Labs has emerged from stealth with $14 million in seed funding and a platform built around a provocative idea: corporate software should be generated from a structured record of business intent, not treated as a pile of source files. Founded by MIT computer science professor Tim Kraska, the company is targeting the coordination problem created by AI coding agents, which Anthropic says now produce up to 80% of the code it sends into production. G5 wants people, agents, policies and architecture to work from the same semantic model.

OpenAI's Codex lands on Dell servers, outside Azure

OpenAI is putting Codex on servers it does not run. Under a partnership announced with Dell, the coding agent — used by more than 4 million developers every week for code review, test coverage, incident response and analysis of large repositories — will run inside customer-owned infrastructure, wired into Dell's data platform for AI. For a company whose commercial architecture has been built…

Nvidia leads AMD by up to 5x in SemiAnalysis's AgentX replay test

SemiAnalysis published AgentX on 24 August, a benchmark that replays recorded coding-agent sessions on production inference stacks rather than firing fixed-length prompts at them. On GLM 5.3 running through open-source SGLang, Nvidia hardware showed up to a fivefold cost-efficiency advantage over AMD at 150 output tokens per second per user. By SemiAnalysis's arithmetic, even if the competing…

NVIDIA hands robot navigation training to a coding agent

NVIDIA has published a workflow that hands most of the labour of training a robot navigation policy to a coding agent. Working with COMPASS, its cross-embodiment mobility framework, a developer names a robot, a scene source and a navigation goal. The agent then checks dependencies, stages assets, runs smoke tests, starts training, investigates failures and compares checkpoints. The human signs…

Anthropic puts a coordinator above Claude Code agents

Anthropic has launched Claude Code Projects, a beta feature that turns Claude Code from a sequence of coding sessions into a persistent project coordinator. Developers can describe a long-running goal in ordinary language, while Claude splits the work across parallel cloud sessions, tracks dependencies, and carries decisions from one stream into the next. The shift matters because software projects rarely end after one prompt or one pull request: they accumulate requirements, exceptions and unfinished work that coding agents must now remember.

UK computer science graduates in coding roles fall from 40% to 28%

The share of British computer science graduates who find professional work as coders or programmers fell from about 40% in earlier years to 28% last year. The share moving into any graduate-level job at all dropped from more than 60% to 50% over two years. The figures come from the Higher Education Statistics Agency, which collected responses from more than 350,000 former students 15 months…

Uber's $1,500 cap and the unproven return on AI tokens

Uber has capped what each employee can spend on each agentic coding tool at $1,500 a month. Microsoft questioned the cost of Claude Code licenses before cancelling their use in its Experiences and Devices division. Duolingo abandoned a plan to factor AI usage into employee performance reviews after staff objected to using the tools for the sake of using them. Three different companies, three…

Students spend four hours making AI slides look worse

Cheating with AI has stopped being a shortcut. A New York University student built a bot that does his calculus homework at the pace of a student who is falling behind, so that the timestamps his professor checks look plausible. A sophomore at UMass Amherst spent four hours degrading a slide deck an OpenAI coding agent had produced, because it looked too professional to pass as his. New York…

OpenAI tells Codex users to cut the rules written for older models

OpenAI is telling developers that the guardrails they wrote for earlier coding models have become the problem. In new guidance on Codex skills and repository instructions, Eric Provancher argues that GPT-6 Astra does better with short skill descriptions, documentation read selectively rather than on every change, and permissions granted in advance — and that the detailed step-by-step…

Meta's new MCP server lets agents set up WhatsApp Business

Meta has released an MCP server that hands the setup of WhatsApp Business messaging to an AI agent. A developer who previously had to move between several Meta tools and services to get a company onto the platform can now describe what is needed to the coding agent of their choice and let it carry out the steps. The server, WhatsApp Business Tools MCP, connects that agent directly to the…

GitHub's HydraFusion cuts cost 67%, matches Opus 5 in one test of three

GitHub has shipped a model router called HydraFusion as a research release in the Copilot CLI. Rather than sending every coding request to one pre-selected model, it decides per request how the work should be done: one model alone, a cheap model that escalates when its answer fails a quality check, or a draft reviewed by a model from a different family. GitHub describes the result as…

Anthropic adds a $35 billion Lambda deal to its pre-IPO compute book

Anthropic has signed a $35 billion deal with Lambda, a week after committing $45 billion to Nscale for a data centre in West Virginia. That is roughly $80 billion in announced compute obligations inside a single week, from a company preparing a public listing. The stated purpose is capacity: meeting rising demand for the Claude model and for Claude Code, the company's coding tool.

Coding agents skip looking at the app when the task gets long

A coding agent is usually evaluated the simple way: hand it a task, a bug report or a test suite, then check whether it fixed the code. A new paper, ProgramDistill , proposes a setup much closer to real work. You have a working reference app. You also have a broken or incomplete version of it. The agent never sees the reference's source. It has…

An AI agent built a playable shooter over 70 autonomous iterations

One of the main problems with coding agents has been obvious to anyone who has tried handing them something bigger than a 50-line function. On a short task everything looks brisk. On a long one the agent tangles itself in its own steps, fixes one thing and breaks another, loses sight of the original requirements, and declares the work finished…

Imagining the poster first lets a coding agent build it in editable layers

Image generators have learned to make beautiful posters. Coding agents have learned to assemble tidy pages in HTML and CSS. Between those two worlds, though, there has always been an awkward gap. An image looks striking, but it is a flattened bitmap: the text inside it often breaks, the layers cannot be pulled apart, the headline cannot be…

Coding agents hit 80% on single tasks, 38% over a whole project

For half a century the industry has run on a simple rule: a person breaks a task into parts, writes code, then fixes and rewrites it as things change. Zhenfeng Cao's paper argues something different: the era in which code was the main carrier of logic is ending . What moves to the front is the AI agent — a system where the LLM does not merely…

Apodex 1.1 gains from agent coordination, but full research runs still fail

Ask a model to write an email, answer a question, even solve a coding problem, and everything looks fine. But the moment a task runs for an hour — reading files, running code, hunting for sources, surviving failures, and finally handing back something you can check — most systems start falling apart. That is what the paper Apodex 1.1: Scaling…

Coding agents solve just 41% of tasks in a new refactoring benchmark

Benchmarks for coding agents follow a familiar arc. First they push the field forward. Then models catch up fast, scores climb, and the metric stops telling systems apart honestly. At the SWE-bench level you can already see it: frontier models pass those tasks more and more often, and some of the unsolved examples turn out to be broken by bad…

8.3 billion simulated users, and the model playing them changes the verdict

Almost every evaluation of AI systems and digital products suffers from the same disease. It tests an averaged abstraction — as if everyone had the same level of experience, the same conversational style, the same tolerance for errors, latency and strange answers. In practice that distorts the picture. A beginner working through a coding task…

SpyRL turns open-ended tasks into a spy hunt with a checkable reward

Large language models have an old problem. They learn well wherever the answer can be checked exactly: math, coding problems, formal puzzles. The answer is either right or it isn't. The machine gets a clean signal and improves. But the moment a task turns open-ended — write a story, summarize a document well, produce a coherent explanation —…

Writing a spec before code lifts coding agents' test pass rate by 21%

AI agents are already decent at fixing code, filling in functions and finding their way around someone else's repository. Ask them to build a program from scratch and the picture changes sharply. Especially when there are no sources at all — just a README and a working binary you can run as a black box. On tasks like these, even the strongest…

A DAG-editing agent matches script baselines at 42.8% lower cost

LLMs already write decent code from a natural-language description. In a production setting that is often not enough. A user asks: "build me a pipeline for data cleaning, filtering, question generation and quality checks." The coding agent answers with a Python script. The script runs. Sometimes it even works. Then the familiar pain starts: it is…

Coding agents fix more bugs when they ask the repo two questions first

Coding agents have an old and very human problem: they start fixing too early . They see a bug report, latch onto familiar words, race through the repository and make a quick edit — in the wrong place, about the wrong thing. Tests fail, steps pile up, tokens burn, and the agent goes in circles. The authors of Know Before Fix propose a simple but…

Agents that can see the tests score 222/222 and skip the library

The AI industry has a favorite trick: post a pretty benchmark score and declare victory. Coding tasks especially. The agent wrote the code, the tests are green, so everything must be fine. But what if that is an illusion? What if the agent passed the exam without building the thing it was asked to build? That is exactly the subject of a paper…

Verifying an AI agent's code is now harder than writing it

There's an old engineering intuition: finding a solution is hard, checking one is easy. For today's coding agents it holds up worse and worse. A model can already produce a plausible patch, page, interface, even a whole repository. What's hard is telling reliably whether the task was actually solved the way the human wanted — and that is the new…

What coding agents transfer across domains is discipline, not code

Coding agents share one weakness: they write code well, but they keep repeating the same mistakes, like an intern who rediscovers every time that running the tests before committing is a good idea. Over the past year researchers have been busy teaching these systems to use memory — to store the moves that worked and the ones that didn't, then…

Training agents on five atomic skills lifts coding scores by 18.7%

The authors propose that we stop training AI agents on composite tasks alone and teach them atomic skills instead — small, checkable, reusable building blocks of software development. Coding agents have a persistent problem: models fit benchmarks well enough, but shift the framing of a task slightly and nothing works. An agent closes…

Auto-generated AGENTS.md files make coding agents worse and costlier

The idea looks obvious: give a coding agent a dedicated file with the repository's rules — how to build the project, how to run the tests, what the conventions are for style and structure — and it will work more confidently and make fewer mistakes. Such files are usually called context files: AGENTS.md, CLAUDE.md and the like. There are already tens of thousands of them in open source, and…

Four agent roles and a real PR review get 72.4% on SWE-bench

LLMs can suggest a chunk of code, explain an error, or write a test. But the moment a task starts to look like real development work — read an issue, find your way around a project, reproduce a bug, produce a patch without breaking everything else — a single general-purpose agent often falls short. In Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering the authors put…

Structured GitHub fix histories lift SWE-Agent by 4.65% on SWE-bench

Once LLMs got good at writing code, a layer of autonomous SWE agents grew up around them: systems that can open a repository, run the tests, localize a bug and prepare a patch. But these agents have an unfortunate habit of working as if they had never seen a bug like this before. Human developers almost never fix hard problems from scratch: they go to GitHub, look for similar issues and PRs,…

Coding agents score 21% when asked to evolve a codebase between releases

Coding agents have gotten noticeably better over the past year: they can find where something broke, edit the code and run the tests. There is an important caveat, though. Most popular benchmarks measure point achievements — fixing one specific bug, or adding a small feature scoped to a single issue. Real development does not work that way. Code lives for years, requirements shift,…

Only 14.5% of agent context files say anything about security

Today, when we write code with AI, we state the task in natural language and the agent inside the IDE plans the steps itself, writes the changes, runs the tests and tries to carry the job through to a result. This is called agentic coding. But it has a weak spot: the agent has to work out fast how the project is put together, what it is conventional to use in it, how to run the build and the…

Making tests fight the patch lifts SWE-bench Verified to 79.4%

LLM-driven bug fixing stopped being exotic a while ago: a model can read code, propose edits, and run the test suite itself. In real repositories, though, the whole thing runs into an awkward detail. There is often no good way to check whether the bug is actually gone. When tests are missing, weak, or simply don't cover the case in question, the system can produce a patch that turns the suite…

DeepCode rebuilds a paper's codebase by managing memory, not context

Over the past year, LLM coding agents really have learned something new: they now handle tests, run commands and get through relatively long sessions. But raise the difficulty — ask an agent to "ship the repository for this paper" — and reality arrives fast. A paper carries a lot of load-bearing detail, scattered across its sections, while the finished project is dozens of files, dependencies,…

Coding agents refactor for readability, not architecture

When AI agents write code, they take on more and more of what used to be purely human work — planning, running tests, even step-by-step refactoring. The authors of Agentic Refactoring: An Empirical Study of AI Coding Agents are the first to look at that practice broadly and in depth: how agents refactor in real open-source projects, whether it pays off, and how their style differs from a human's.

BigCodeArena scores coding models by running their code, not reading it

Judging code generation by how clean the comments look is like judging a car from the brochure. In real life what matters is whether it starts, whether it brakes in time, and whether it is pleasant to use. The authors of BigCodeArena take exactly that practical view: their open platform compares large language model (LLM) solutions not by string overlap but by whether they run, how they…

Natural-language memory beats fine-tuning on long agent tasks

Large language models do well on short reasoning and coding benchmarks. But real work stretches over dozens or hundreds of steps, demands switching between applications, careful context management and the ability to catch your own mistakes. The core problem is that at test time most agents stay static: they accumulate no experience and get no better from one attempt to the next. The authors of…

Ranking code by perplexity compresses context 5.6× without hurting accuracy

Coding LLMs can already complete, explain and fix code, but real projects make them read thousands of lines. Long context windows help, yet they cost time and money — and they paradoxically hurt accuracy: the model starts drowning in detail and misses hidden dependencies between files and functions. Generic text compressors throw away phrases and tokens with no regard for code structure.…

Top coding agents solve under a quarter of SWE-Bench Pro tasks

Over the past couple of years, agents built on large language models have settled into everyday development: they read repositories, fix bugs, propose patches and run tests. On the classic SWE-Bench Verified, the top systems clear more than 70% of tasks on the first attempt. The trouble is that progress like this creates a false sense of readiness for real industry work. Out there, tasks…