Train frontier ideas, not infrastructure teams.
Take a novel architecture from a local experiment to a frontier run without inheriting one vendor’s software stack. SF Tensor optimizes, verifies, places and operates training across the best available compute.
Input
Optimize
Generate fast kernels
Scale
1 to 10,000 GPUs
Recover
Resume automatically after node failure
Outcome
01. Average fleet utilization
Measured across a set of AI labs running SF Tensor.
02. GPUs, one workflow
From local research to a production pre-training fleet.
03. Failure recovery
Resume from the last checkpoint after node failure, with no human in the loop.
Frontier runs should not require frontier-lab headcount.
The scarce thing is the architecture and the people inventing it. We make compute supply fungible by retargeting workloads across vendors, compiling kernels for the actual topology and operating the fleet, so a small team can move like a large lab.
01
Bring the training repo
Keep the framework and model code your researchers already use.
02
Compile for the fleet
Search, generate and formally verify optimized kernels for each target.
03
Scale across supply
Place the job on the best available mix of hardware, providers and regions.
04
Run the experiment
Resume automatically from the last checkpoint after node failure, with no human in the loop.
The systems team behind every training run.
Use SF Tensor as the compiler, performance, distributed systems and fleet operations team that turns research code into a reliable frontier run.
Novel architectures welcome
Bring unusual layers, kernels, communication patterns and research frameworks.
Automatic kernel optimization
Generate workload-specific kernels instead of waiting months for hand-tuning.
Cross-vendor portability
Retarget training as new accelerators and lower-cost supply enter the market.
Training-native storage
Keep data and checkpoints close to workers with shared, cluster-local caching.
Resilient fleet operations
Handle placement, observability, checkpoints, node failures and run recovery.
Give every research direction a credible path to scale.
Foundation models
Pre-train any foundation model from scratch: dense, MoE, multimodal, domain-specific or something no one has named yet.
Arbitrary architectures
Not a fixed menu of blessed model types. Any attention scheme, state-space model, memory system, routing or communication design, including ones that don't exist yet.
Alternative silicon
Use emerging accelerators when they are the best technical or economic fit.
Your architecture is the bet. We make sure the infrastructure never is.
Bring the training repo. We optimize, place and operate the run from 1 GPU to 10,000.
Train the models only you can build. One stack for enterprise post-training and frontier pre-training.
All Systems Operational© 2026 San Francisco Tensor Company
