Skip to content
Plan a training run

What we learned topping NVIDIA's SOL-ExecBench

We're #1 on SOL-ExecBench across all four tracks and learned a lot of bitter lessons about frontier models, reward hacks and why the benchmark is running out of time

What we learned topping NVIDIA's SOL-ExecBench

SOL-ExecBench is NVIDIA's official kernel optimization benchmark for the B200, made up of 235 kernels extracted from real production workflows that you need to optimize.

Since Amir first reached #1 on SOL-ExecBench on the 25th of July, things have changed quite a bit and we've learned a lot. This post covers how the race played out, what we saw from more than a dozen models across tens of thousands of agent rollouts, and why we think the benchmark is close to saturated.

Results

As of September 25, across 31,466 submissions from 583 teams, we are:

  1. #1 overall, with an average SOL score of 0.7992 across all 235 kernels
  2. #1 in every track, across L1 single ops, L2 fused ops, Quantization (FP8/NVFP4) and FlashInfer-Bench
  3. 3.03x median speedup per kernel over NVIDIA's optimized reference
#TeamScoreL1L2QuantFlashInfer
1SF Tensor0.79920.79040.77890.88110.7916
2RIAC as well0.79740.78860.77780.87630.7904
3Geometric0.76520.75040.73370.87050.7845
4ac4k0.76010.75630.73400.83080.7670
5Hyra0.73590.74650.72370.74100.7297
6doubleAI0.72410.72650.70030.78840.7085
7Recursive0.67570.70300.65180.68890.6357
8Cursor0.49000.54040.40070.48000.6022

Update September 26: we passed an average SOL score of 0.8.

It's been becoming an increasingly close race. Over the last few weeks we've seen the gap between 1st and 2nd place shrink at an increasingly fast pace.

The problem is that when two teams with different agents, models, and approaches land within 0.01 SOL of each other on a kernel, the truth is that said kernel is probably close to its practical ceiling. In mid-July that was true for 39% of kernels. It's now true for 65%, and on 34% of kernels the top three teams are within 0.01 of each other.

Independent teams are landing on the same answer
Share of 235 kernels where other teams are within 0.01 SOL of the #1 score.
#2 within 0.01#2 and #3 within 0.01
By September 25, #2 was close on 152 kernels and both #2 and #3 were close on 80.

The benchmark is saturating

The problem is that when a majority of kernels in a benchmark are already at their practical ceiling, the shrinking gap between the top teams stops saying much about the teams and starts being about the benchmark.

Convergence is strongest when scores are highest. 32 kernels have a best-known score of 0.95 or higher, and on 21 of them, three or more teams are within 0.01 of the top. The remaining headroom is concentrated in a long tail of harder kernels, and as we discuss below, some of the headroom is an artifact of how the benchmark scores and checks results rather than real room to optimize.

Where the ceiling is
Best-known SOL score for each kernel, sorted highest to lowest. 1.0 is the theoretical hardware bound.

Loading kernel scores…

3+ teams within 0.01 of #1Fewer than 3 teams
32 of 235 kernels have reached 0.95 SOL or higher. Blue bars have at least three teams within 0.01 of the leader.

The problem with a saturated benchmark is that the runtime of kernels depend heavily on its environment, causing fluctuations upward of 5% on "the same" hardware. These factors become increasingly relevant and make a #1 score less significant as an evaluation of the tools used.

Model Behavior

We've had tens of thousands of agent rollouts across over a dozen models, spanning millions of agent steps and well over 100B tokens and observed some interesting behavior across them

GPT-6 Astra

Astra is an exceptionally strong model, but requires very precise instructions as it otherwise makes some very naive mistakes, such as starting from a random kernel in the library instead of starting from the fastest or other behavior that causes severe performance regression.

GPT-5.6 Sol

It will start off and then quickly come up with a big bet on what it could do to create a significantly faster kernel, but when that kernel does not compile or return correct results very quickly, sometimes on first try, it will fallback to an approach that it knows to work and that produces correct results, even if very slow and will then get stuck in that local minima and spend ages trying to hyper-tune or optimize the specific kernel by doing things such as making minor changes to memory layouts, even in cases where the kernel and approach is clearly suboptimal, but will continue with this behavior instead of attempting to take the next obvious big step, such as fusing two kernels or trying to resolve the correctness problems of its earlier abandoned big step.

Task: Fuse 7 ops into a MegaKernel

Instead of implementing a MegaKernel the way we would expect, it merged the 7 separate kernels into a single device-side function, with one kernel running to completion before the next one ran, all device-side. To do that, it had to implement an atomic-based synchronization where as each thread arrived, it performed an atomic increment, creating a global barrier, which does not exist in CUDA for good reasons (barring cooperative cooperative groups, which the agent was not allowed to use): if we launch a kernel with 10,000 blocks, the GPU might only have enough physical SMs to run 100 blocks at the same time. After those blocks run and hit the global barrier, they will have to wait for the remaining blocks to reach it, causing the GPU to deadlock instantly. Even better, instead of just producing this monstrosity for itself, it decided that this wonderful pattern was so valuable that it must be recorded for future agents.

cuda__device__ unsigned int bar_count = 0;   // must be 0 at launch

__device__ void grid_barrier_once() {
    __syncthreads();
    if (threadIdx.x == 0) {
        unsigned int nblocks = gridDim.x * gridDim.y * gridDim.z;
        __threadfence();                                  // publish writes
        atomicAdd(&bar_count, 1);
        while (atomicAdd(&bar_count, 0) < nblocks) { }    // spin
        __threadfence();
    }
    __syncthreads();
}

Opus 5

While many models ran into cases where they were unable to make the kernels any faster, Opus 5 likes to eventually conclude that it's not him who can't get any faster, it's simply that there is no way to make this kernel any faster and states so in a matter-of-factly attitude.

Another interesting behavior of Opus 5 is that unlike other models which generally make their first submissions relatively quickly, allowing them to iterate on the kernels based on the results from the grader, Opus 5 takes much longer to make its first submission.

DeepSeek V4 Pro

The model was extremely fast at reaching almost precisely the performance of first place on approximately a third of the tasks, significantly faster than substantially smarter models took to reach that point, behavior that is consistent with a model that was post-trained or contaminated on those tasks. It did not perform well at all on the remaining tasks or held out tasks from our own evals.

Interesting model behaviors & reward hacks

Why bother computing what was asked?

Given the lax numeric bounds, an agent realized that on an MoE task that asked for it to route to the top-10 experts, run the experts and then perform the expert aggregation, the probability distribution was skewed heavily enough that if it routed only to the top-2 experts, ran them and then performed the aggregation, the result was numerically close enough to the top-10 routing that it passed the tests, even though it only ran 20% of the required experts. Despite this, the agent did not make it into the top results because the MoE was extremely sparse and therefore memory-bound; it just went from bad to slightly less bad despite blatantly cheating.

ZipGate

At some point, an agent wanted to copy a CUTLASS reference implementation into its workspace to be able to work from that. The agents have a bash tool that they can and regularly do call that lets them perform arbitrary bash commands within their sandbox, so one would assume that the model just performed a simple copy. Unfortunately, that was not quite the case. Looking into the model's reasoning trace, we saw that it wanted to copy the content but was afraid that the file was too large to read and output from memory in one go, so it decided to compress the reference implementation using zip in its source directory, then read the contents of the zip file into its memory before writing it into a file in its own workspace from memory and decompressing the file.

Why not just change the grader's reference implementation?

An agent realized that it could set environment variables during its imports to modify the CUBLAS behavior of the grader. Because the grader and the submitted kernel run in the same process, this let the agent change the grader's reference implementation to make it run its matmuls at a lower precision to match the agent's low-precision matmuls. This brought the result back within the numeric bounds, not by making itself more accurate but by reducing the accuracy of the reference.

BoomerCores

For a while, when asked simply to optimize a kernel for Blackwell using the latest hardware features without being provided any further information and despite running in a sandbox with a Blackwell GPU, most models defaulted to using the old Ampere MMA operations, which produced valid results but did not utilize Blackwell's 5th-generation tensor cores, substantially limiting the performance the models could reach. Resolving this simply required providing the agents with some instructions to use tcgen05.

Problems with SOL-ExecBench

To start off, SOL-ExecBench is an exceptional benchmark that came at a time when alternatives such as KernelBench were severely limited, had a whole host of other problems and provided no easy way to compare your results with others.

That said, it's been 6 months, and with the benefit of hindsight and lots of runs, we now know that there are certain flaws.

Numerical Correctness Bounds

The numerical correctness bounds are substantially too lax on certain problems, such as the MoE routing kernel described above, where you can skip 80% of experts and still pass. On other problems, they are too tight, allowing for barely any deviation from even the accumulation order of the reference implementation and making it nearly impossible to achieve substantial optimization.

SOL-Score over 1

The idea of the SOL score is to grade a result between 0 and 1, with 0.5 being their optimized baseline and 1.0 being the theoretical hardware lower bound. In one MoE task, the hardware lower bound assumes that all experts must be loaded, but in some cases the routing is not uniformly distributed and therefore, in some batches, not all experts have to be loaded, reducing the amount of data movement that is necessary on the lower bound, letting you "beat the Speed of Light".

Grader and Candidate run in the same process

Despite running arbitrary, unvalidated kernels, the grader, reference implementation and candidate all run in the same process, allowing an adversarial candidate kernel to do things such as modify global environment variables, which have an impact on PyTorch and CUBLAS behavior, or going even further, running arbitrary host-side code to "hijack" the grader.

What's next

At SF Tensor, our kernel optimization work covers a significantly wider range than what SOL-ExecBench covers, and being able to evaluate models in a way that provides significant meaning is essential to that effort, especially as we've started post-training our own models.

Toward that end, we've been working on KernelWorld, our internal suite of environments and test cases for GPU kernels. It already comprises 22k unique environments and over 484k test cases covering numerous hardware types per vendor (e.g., A100, H100, B200, GB300), numerous vendors (NVIDIA, AMD, TPU, Trainium and some others) as well as multi-GPU and multi-node evaluations.

From this, we've been holding out a slice of environments and test cases for evaluation, which we've been using internally to benchmark models.

Results on KernelWorld-Eval v0.1
Kernel Eval score by model. Higher is better.