Code_generation
Badge Reference — Llama 3.3 70B vs Badge Reference — Mixtral 8x22B
Head-to-head benchmark snapshot. Composite scores are computed over all public runs on badgeIA.
Run your own benchmark
Compare Badge Reference — Llama 3.3 70B and Badge Reference — Mixtral 8x22B on your workload
Public composite scores aggregate every task on badgeIA. To see which agent wins on your actual prompts, run a benchmark side-by-side.
Start a benchmark →