Code Generation · Language Models · Open Models

AlphaCode vs Code Llama: Million-Sample Contest Search or Open Weights?

AlphaCode converts up to a million sampled candidates into a top-54.3% Codeforces finish; Code Llama ships open 7B-70B weights that hit 67% HumanEval: one is search, the other is infrastructure.

AlphaCode vs Code Llama: Million-Sample Contest Search or Open Weights?

The same headline, two different machines

Both papers promise “an AI that writes programs,” and both delivered something real. AlphaCode, from DeepMind in 2022, was the first system to post a roughly median result in simulated real programming competitions: a top 54.3% average ranking against fields of more than 5,000 participants. Code Llama, from Meta AI in August 2023, was the best openly available code model family at its release, scoring up to 67% on HumanEval and 65% on MBPP under a license allowing commercial use.

But the two systems answer different questions. AlphaCode asks: given a contest problem, a strict submission limit, and a model whose individual guesses are usually wrong, how much correctness can you buy back with brute-force sampling and behavioral selection? Code Llama asks: how good can a single forward pass get if you keep training a strong open base model on code, and does giving the weights away change what the ecosystem builds?

One is an inference-time system built around a model. The other is a model built to be deployed. Almost every confused take on this pair traces back to mixing those two categories up, and it is exactly why the headline numbers, 54.3% and 67%, cannot be lined up side by side.

How AlphaCode’s pipeline works

The AlphaCode paper’s own framing is that competitive programming is a different sport from the benchmarks of its era. HumanEval-style tasks hand you a short, self-contained function description. A Codeforces problem is a multi-paragraph story you must first decode into a formal specification, then invent an algorithm with the right time complexity, implement it, and survive hidden tests that include adversarial edge cases. A solution that is 95% right scores zero. There is no partial credit, which is precisely the environment where a language model’s “plausible but wrong” failure mode is fatal.

The system attacks this with a pipeline rather than a bigger model. Training data scraped naively from competitive programming is noisy: solutions pass visible tests but fail hidden ones, so the team built a dataset cleaned with extra generated test cases. Architecturally they chose an asymmetric encoder-decoder with a large encoder and a smaller decoder, plus multi-query attention, specifically so that drawing orders of magnitude more samples per problem stays affordable.

At inference time the numbers get extreme. AlphaCode generates up to a million candidate programs per problem. Roughly 99% of them fail the example tests printed in the problem statement and are discarded. The survivors are clustered by how they behave on model-generated inputs, and the system submits one representative per cluster, respecting the ten submissions a contest allows. The top 54.3% ranking is the output of that whole machine (model, filter, and cluster), not of the model alone.

The honest implication: a single sample from AlphaCode is very unlikely to solve a hard contest problem. The contribution is a demonstration that weak per-sample success, multiplied by massive sampling and disciplined selection, converts into a real contest result. It is a search result wearing a generation model’s clothes.

How Code Llama works

Code Llama starts from the opposite corner. It does not train a code model from scratch; it continues training Llama 2 on code-heavy data and adds task-specific stages, so the code models inherit the base model’s language reasoning instead of paying for it twice. The release spans 7B, 13B, 34B, and 70B sizes in three flavors: the base model for general generation, Code Llama - Python for Python-first teams, and Code Llama - Instruct for instruction-following assistant use.

Two engineering choices target real developer workflows rather than benchmarks. First, infilling: the 7B, 13B, and 70B base and Instruct variants can fill a missing span using the code both before and after the gap, which is what an IDE autocomplete actually needs because nobody writes code strictly left to right. Second, context: all models train on 16k-token sequences and show improvements on inputs up to 100k tokens, so you can feed a model a meaningful slice of a real repository instead of one isolated function.

The headline scores, up to 67% on HumanEval and 65% on MBPP, came from the largest, most specialized configurations at release. The sharper single data point in the paper is a size-versus-specialization comparison: Code Llama - Python 7B outperforms its own base model, general-purpose Llama 2 70B, on both benchmarks. A model one-tenth the size beats the general model purely through code specialization, and all variants beat every other publicly available model on MultiPL-E, the multi-language coding benchmark. Everything ships under a license permitting research and commercial use, which is the part that mattered to the ecosystem: this is the model people actually built on.

Key numbers

MeasurementAlphaCodeCode LlamaSetting and sourceSame harness?
Codeforces contest resulttop 54.3% average ranking, fields of 5,000+not evaluated on itAlphaCode, simulated recent contests, 10 submissions allowedNo Code Llama number in this set
HumanEvalnot reported in our sourceup to 67%Code Llama, largest configs; best open scores at releaseNo AlphaCode number in this set
MBPPnot reported in our sourceup to 65%Code Llama, same releaseNo AlphaCode number in this set
Samples per taskup to 1,000,000 candidates, ~99% cut by example tests, clustered down to 10 submissions1 completion per callAlphaCode pipeline vs Code Llama deploymentn/a
Size ladderasymmetric encoder-decoder (large encoder, small decoder), multi-query attention for cheap sampling7B / 13B / 34B / 70Bboth papersn/a
Specialization proofnot reportedCode Llama - Python 7B beats Llama 2 70B on both HumanEval and MBPPCode Llama papern/a
Context lengthnot the bottleneck for single problems16k-token training, gains up to 100k tokensCode Llama papern/a
Multi-languagenot reportedall variants best public model on MultiPL-ECode Llama papern/a

Read the table for what it refuses to give you: there is no row where both systems are measured on the same task under the same conditions. Neither paper evaluates on the other’s benchmark: AlphaCode reports contest rankings, Code Llama reports pass-style scores on short function benchmarks. Any claim of the form “AlphaCode gets X%, Code Llama gets Y%, therefore one is better” is comparing a ranking percentile earned under a submission cap against a single-shot pass rate on a different task family.

Why the headline numbers don’t compare

Three mismatches sit between 54.3% and 67%.

Different tasks. Codeforces problems require decoding a story into a formal spec, choosing an algorithm, and passing adversarial hidden tests; a 95% correct solution scores zero. HumanEval and MBPP ask for short, self-contained functions where the prompt mostly describes the implementation. The first tests algorithm invention under a judge; the second tests local implementation fluency.

Different metric semantics. A top 54.3% average ranking means the system, as a whole, finished ahead of about 46% of a 5,000+ person field when allowed ten submissions. A 67% HumanEval score means a single generation passes the unit tests for about two-thirds of problems. One number is a tournament placing bought with a submission budget; the other is a per-shot success rate. They don’t even share a denominator.

Different sampling budgets. AlphaCode’s result is explicitly the product of up to a million candidates, a ~99% cut at the example-test filter, and clustering down to ten submissions. Code Llama’s 67% is one forward pass. Run the comparison at equal sample counts and neither headline survives: a single-shot AlphaCode would be near-useless on hard problems, while Code Llama’s per-sample rate was never measured under a million-sample search. The papers are simply not in the same experimental design.

When to use which

The contest setting is where AlphaCode’s philosophy earns its keep. If your problem has an automatic judge, hidden tests, and an allowance for many attempts (competitive programming, some synthesis-and-verify research loops), then sampling broadly and selecting behaviorally is a proven path from a modest model to results that look impossible single-shot. The modern descendants of this bet are the agentic search-and-harness approaches that generate, execute, and filter at scale, a lineage that continues on this site in code as an agent harness.

Everything else points at Code Llama’s answer. IDE autocomplete, infilling into the middle of a file, repo-level context up to 100k tokens, self-hosting, fine-tuning for an internal codebase, shipping a commercial product on top of open weights: all of these need one good completion per call, not a million. Add a permissive license and three size tiers, and it is obvious why Code Llama became infrastructure while AlphaCode remained a landmark research result. Notably, both papers predate the agent era, and neither claims to handle multi-file repository work: that frontier is being fought by different systems entirely.

Limits and open questions

No shared benchmark exists in this pair, and per our sourcing rules this page won’t invent one. Both papers also evaluate in ways that flatters them: contest simulations for AlphaCode, saturated short-function benchmarks for Code Llama, and the Code Llama page itself warns that HumanEval and MBPP say little about debugging a real repository or maintaining a large codebase. The headline 67%/65% belongs to the largest configurations; the 7B a laptop can run scores meaningfully lower. AlphaCode’s pipeline, meanwhile, is a research artifact: the compute bill for a million samples per problem is not a product strategy, it is a scientific statement about where correctness can come from. The open question both papers leave behind is the same one: when does search stop being the expensive way to look smart, and when does a strong enough single model make it obsolete?

FAQ

Is AlphaCode better than Code Llama?

They are not directly comparable, because they were built for different jobs. AlphaCode is an inference-time system that converts up to a million sampled candidates per problem, filtered and clustered down to ten submissions, into a top 54.3% average ranking on simulated Codeforces contests. Code Llama is a family of open models scoring up to 67% on HumanEval and 65% on MBPP in a single forward pass. One is a search pipeline around a model; the other is deployable weights.

Why does AlphaCode report a Codeforces ranking while Code Llama reports HumanEval percentages?

Because they answer different questions. Codeforces ranking percentile measures tournament performance under a ten-submission cap against thousands of humans, with adversarial hidden tests and zero partial credit. HumanEval and MBPP percentages measure how often a single generation passes unit tests on short, self-contained functions. Neither paper reports the other’s metric, so no same-harness comparison exists.

Can Code Llama solve competitive programming problems?

The Code Llama paper does not evaluate on contest problems, so there is no sourced number for this. What is known is that Code Llama dominates benchmarks of short function-writing tasks, while contest problems demand algorithm selection under hidden adversarial tests, a different skill profile. Systems built in AlphaCode’s style, which sample massively and select behaviorally, are the ones designed for that setting.

How many samples does AlphaCode use per problem?

Up to one million candidate programs per problem. Roughly 99% of them are discarded for failing the example tests in the problem statement, the survivors are clustered by behavior on model-generated inputs, and the system submits one representative per cluster, respecting the ten-submission contest limit. The top 54.3% ranking is the output of the entire pipeline, not of a single generation.

Which one can I actually run today?

Code Llama. Its weights are open under a license permitting research and commercial use, in 7B, 13B, 34B, and 70B sizes, with infilling support on the 7B, 13B, and 70B base and Instruct variants. AlphaCode was released as a research result; its model is one component of a pipeline whose per-problem compute cost (up to a million samples) is not something you deploy.

What is the legacy of each approach?

AlphaCode’s legacy is the idea that generation plus large-scale sampling and behavioral selection can beat a single smarter generation: the seed of today’s agentic coding harnesses. Code Llama’s legacy is the open-weights code model as ecosystem infrastructure: a permissively licensed base that everyone could fine-tune, quantify, and ship, proving that code specialization lets a 7B Python model beat a 70B general one.