i
DATAIST
Review · 2026-01-02

Reasoning models now pass all three CFA levels, but ethics still trips them up

Reasoning models now pass all three CFA levels, but ethics still trips them up

In finance, the CFA (Chartered Financial Analyst) exams are a marathon run over three distances. Level I tests the fundamentals and whether you can keep the terminology straight. Level II puts you in front of cases where formulas and logic have to be applied in context. Level III asks for more than correct arithmetic: a coherent professional answer — how to build a portfolio, how to assess risk, what to do when a situation is ambiguous, and where the ethical line falls.

Until recently, large language models (LLMs) looked unconvincing on these exams: they could show off facts, but fell apart on applied problems and especially on questions that demand reasoning. The authors of Reasoning Models Ace the CFA Exams show that this picture is changing fast. Today's reasoning models clear all three levels with unexpected confidence — and they do it across a large set of practice exams.

Examples of how CFA practice questions are built, level by level: from short multiple-choice items to cases with long context.

Why the CFA is such a revealing stress test for AI

Finance exams are useful because they mix everything at once. You have to calculate, compare, interpret text, recall definitions, catch nuances, and sometimes pick the least wrong option in an ethical dilemma. That is closer to an analyst's real job than a standard question-and-answer benchmark. Which is why the CFA has long been a convenient yardstick: if a model passes it consistently, progress in reasoning is not cosmetic.

There is an important catch, though. Earlier studies often leaned on outdated question sets. The CFA curriculum is updated, its emphases shift, new formats appear. The authors of the new paper try to avoid that trap and build their dataset from current materials.

How the test set was built and scored

For the experiment the authors take 980 questions from several practice exams: three Level I versions, two Level II versions and three Level III versions. The sources are the official CFA Institute Practice Pack (for Levels I–II) and AnalystPrep (for Level III). That choice matters: it lowers the chance the model has simply seen these items before, and raises the odds that what is being measured is the ability to reason rather than to recall.

Two answering strategies are tested. The first gives no explanation, only the final choice. The second asks the model to reason step by step (chain of thought). For multiple-choice items scoring is simple: percentage correct. For the constructed responses at Level III, an automatic grader built on o4-mini compares the model's answer against a reference answer and a rubric. The authors are candid about the bias this introduces: graders like this sometimes favor longer answers and can under-penalize subtle errors.

What the models scored

The central question of the paper is the comparison across generations. Older models really did fail the levels often, especially II and III. But almost every current reasoning model passes all three.

Among the leaders on overall performance the authors single out Gemini 3.0 Pro, Gemini 2.5 Pro, GPT-5, Grok 4, Claude Opus 4.1 and DeepSeek-V3.1. The headline numbers look almost implausible for an exam of this class: Gemini 3.0 Pro scores 97.6% on Level I, and the best Level II result belongs to GPT-5 at 94.3%.

One curious observation: chain of thought helps older models noticeably, but on strong reasoning models the effect is mixed. On multiple-choice items accuracy sometimes even dips slightly, while on constructed responses the reasoning mode more often improves quality.

Where models still get it wrong — and why that matters

The authors work through typical failures, and they make it clear that even at high overall accuracy the weak spots stay recognizably human: misreading the setup, confusing statements that look alike, applying a rule incorrectly, over-simplifying, or plain slips in the arithmetic.

An example where the model misapplies ethical standards to a specific situation — one of the stickiest question types.
A calculation error: the model plugs in the wrong base values and arrives at the wrong financial result.

Ethics gets called out separately as one of the hardest areas. That makes sense: those questions hang on context and fine differences in wording rather than on formulas.

What this changes, and what the authors want next

The work sets a new reference point: "can an LLM pass the CFA?" is now mostly yes, at least in practice-exam form. But the authors do not want the conclusion to be that AI has become a financial analyst. They carefully record the limits instead: Level III does not come from an official source, and the constructed responses are graded by an automatic judge.

Put in practical terms, the CFA turns out to be less an exam of knowledge than an instrument of fine-grained diagnosis: where the model reasons, where it guesses, where it writes well but is wrong on substance. The next step is stricter evaluation with human examiners involved.

AI paper breakdowns

Every day we read the new AI papers and retell what matters in plain language — no hype, no filler. If you want to see where AI agents are heading before everyone else, subscribe.

New breakdowns every day.

On Telegram