SOL-ExecBench is NVIDIA's official kernel optimization benchmark for the B200, made up of 235 kernels extracted from real production workflows that you need to optimize.
Since Amir first reached #1 on SOL-ExecBench on the 25th of July, things have changed quite a bit and we've learned a lot. This post covers how the race played out, what we saw from more than a dozen models across tens of thousands of agent rollouts, and why we think the benchmark is close to saturated.
Results
As of September 25, across 31,466 submissions from 583 teams, we are:
- #1 overall, with an average SOL score of 0.7992 across all 235 kernels
- #1 in every track, across L1 single ops, L2 fused ops, Quantization (FP8/NVFP4) and FlashInfer-Bench
- 3.03x median speedup per kernel over NVIDIA's optimized reference
| # | Team | Score | L1 | L2 | Quant | FlashInfer |
|---|---|---|---|---|---|---|
| 1 | SF Tensor | 0.7992 | 0.7904 | 0.7789 | 0.8811 | 0.7916 |
| 2 | RIAC as well | 0.7974 | 0.7886 | 0.7778 | 0.8763 | 0.7904 |
| 3 | Geometric | 0.7652 | 0.7504 | 0.7337 | 0.8705 | 0.7845 |
| 4 | ac4k | 0.7601 | 0.7563 | 0.7340 | 0.8308 | 0.7670 |
| 5 | Hyra | 0.7359 | 0.7465 | 0.7237 | 0.7410 | 0.7297 |
| 6 | doubleAI | 0.7241 | 0.7265 | 0.7003 | 0.7884 | 0.7085 |
| 7 | Recursive | 0.6757 | 0.7030 | 0.6518 | 0.6889 | 0.6357 |
| 8 | Cursor | 0.4900 | 0.5404 | 0.4007 | 0.4800 | 0.6022 |
Update September 26: we passed an average SOL score of 0.8.
It's been becoming an increasingly close race. Over the last few weeks we've seen the gap between 1st and 2nd place shrink at an increasingly fast pace.
The problem is that when two teams with different agents, models, and approaches land within 0.01 SOL of each other on a kernel, the truth is that said kernel is probably close to its practical ceiling. In mid-July that was true for 39% of kernels. It's now true for 65%, and on 34% of kernels the top three teams are within 0.01 of each other.
The benchmark is saturating
The problem is that when a majority of kernels in a benchmark are already at their practical ceiling, the shrinking gap between the top teams stops saying much about the teams and starts being about the benchmark.
Convergence is strongest when scores are highest. 32 kernels have a best-known score of 0.95 or higher, and on 21 of them, three or more teams are within 0.01 of the top. The remaining headroom is concentrated in a long tail of harder kernels, and as we discuss below, some of the headroom is an artifact of how the benchmark scores and checks results rather than real room to optimize.
The problem with a saturated benchmark is that the runtime of kernels depend heavily on its environment, causing fluctuations upward of 5% on "the same" hardware. These factors become increasingly relevant and make a #1 score less significant as an evaluation of the tools used.
Model Behavior
We've had tens of thousands of agent rollouts across over a dozen models, spanning millions of agent steps and well over 100B tokens and observed some interesting behavior across them
GPT-6 Astra
Astra is an exceptionally strong model, but requires very precise instructions as it otherwise makes some very naive mistakes, such as starting from a random kernel in the library instead of starting from the fastest or other behavior that causes severe performance regression.
GPT-5.6 Sol
It will start off and then quickly come up with a big bet on what it could do to create a significantly faster kernel, but when that kernel does not compile or return correct results very quickly, sometimes on first try, it will fallback to an approach that it knows to work and that produces correct results, even if very slow and will then get stuck in that local minima and spend ages trying to hyper-tune or optimize the specific kernel by doing things such as making minor changes to memory layouts, even in cases where the kernel and approach is clearly suboptimal, but will continue with this behavior instead of attempting to take the next obvious big step, such as fusing two kernels or trying to resolve the correctness problems of its earlier abandoned big step.
Opus 5
While many models ran into cases where they were unable to make the kernels any faster, Opus 5 likes to eventually conclude that it's not him who can't get any faster, it's simply that there is no way to make this kernel any faster and states so in a matter-of-factly attitude.
Another interesting behavior of Opus 5 is that unlike other models which generally make their first submissions relatively quickly, allowing them to iterate on the kernels based on the results from the grader, Opus 5 takes much longer to make its first submission.
DeepSeek V4 Pro
The model was extremely fast at reaching almost precisely the performance of first place on approximately a third of the tasks, significantly faster than substantially smarter models took to reach that point, behavior that is consistent with a model that was post-trained or contaminated on those tasks. It did not perform well at all on the remaining tasks or held out tasks from our own evals.
Interesting model behaviors & reward hacks
Why bother computing what was asked?
Given the lax numeric bounds, an agent realized that on an MoE task that asked for it to route to the top-10 experts, run the experts and then perform the expert aggregation, the probability distribution was skewed heavily enough that if it routed only to the top-2 experts, ran them and then performed the aggregation, the result was numerically close enough to the top-10 routing that it passed the tests, even though it only ran 20% of the required experts. Despite this, the agent did not make it into the top results because the MoE was extremely sparse and therefore memory-bound; it just went from bad to slightly less bad despite blatantly cheating.
ZipGate
At some point, an agent wanted to copy a CUTLASS reference implementation into its workspace to be able to work from that. The agents have a bash tool that they can and regularly do call that lets them perform arbitrary bash commands within their sandbox, so one would assume that the model just performed a simple copy. Unfortunately, that was not quite the case. Looking into the model's reasoning trace, we saw that it wanted to copy the content but was afraid that the file was too large to read and output from memory in one go, so it decided to compress the reference implementation using zip in its source directory, then read the contents of the zip file into its memory before writing it into a file in its own workspace from memory and decompressing the file.
Why not just change the grader's reference implementation?
An agent realized that it could set environment variables during its imports to modify the CUBLAS behavior of the grader. Because the grader and the submitted kernel run in the same process, this let the agent change the grader's reference implementation to make it run its matmuls at a lower precision to match the agent's low-precision matmuls. This brought the result back within the numeric bounds, not by making itself more accurate but by reducing the accuracy of the reference.
BoomerCores
For a while, when asked simply to optimize a kernel for Blackwell using the latest hardware features without being provided any further information and despite running in a sandbox with a Blackwell GPU, most models defaulted to using the old Ampere MMA operations, which produced valid results but did not utilize Blackwell's 5th-generation tensor cores, substantially limiting the performance the models could reach. Resolving this simply required providing the agents with some instructions to use tcgen05.
Problems with SOL-ExecBench
To start off, SOL-ExecBench is an exceptional benchmark that came at a time when alternatives such as KernelBench were severely limited, had a whole host of other problems and provided no easy way to compare your results with others.
That said, it's been 6 months, and with the benefit of hindsight and lots of runs, we now know that there are certain flaws.
Numerical Correctness Bounds
The numerical correctness bounds are substantially too lax on certain problems, such as the MoE routing kernel described above, where you can skip 80% of experts and still pass. On other problems, they are too tight, allowing for barely any deviation from even the accumulation order of the reference implementation and making it nearly impossible to achieve substantial optimization.
SOL-Score over 1
The idea of the SOL score is to grade a result between 0 and 1, with 0.5 being their optimized baseline and 1.0 being the theoretical hardware lower bound. In one MoE task, the hardware lower bound assumes that all experts must be loaded, but in some cases the routing is not uniformly distributed and therefore, in some batches, not all experts have to be loaded, reducing the amount of data movement that is necessary on the lower bound, letting you "beat the Speed of Light".
Grader and Candidate run in the same process
Despite running arbitrary, unvalidated kernels, the grader, reference implementation and candidate all run in the same process, allowing an adversarial candidate kernel to do things such as modify global environment variables, which have an impact on PyTorch and CUBLAS behavior, or going even further, running arbitrary host-side code to "hijack" the grader.
What's next
At SF Tensor, our kernel optimization work covers a significantly wider range than what SOL-ExecBench covers, and being able to evaluate models in a way that provides significant meaning is essential to that effort, especially as we've started post-training our own models.
Toward that end, we've been working on KernelWorld, our internal suite of environments and test cases for GPU kernels. It already comprises 22k unique environments and over 484k test cases covering numerous hardware types per vendor (e.g., A100, H100, B200, GB300), numerous vendors (NVIDIA, AMD, TPU, Trainium and some others) as well as multi-GPU and multi-node evaluations.
From this, we've been holding out a slice of environments and test cases for evaluation, which we've been using internally to benchmark models.

