Breakdowns of work on autonomous agents: planning, tool calls, multi-agent systems and why they fall apart on long tasks.
7 articlesTopics
The same posts, sorted by subject instead of date.
Models that write, read and fix code — from autocomplete to working through a task in a repository on their own.
1 articleHow model quality is measured and why benchmarks keep lying.
1 articleGetting more out of a model: prompting, picking the right mode, and the mistakes people make using it.
1 article