Plan a training run

Train frontier ideas, not infrastructure teams.

Take a novel architecture from a local experiment to a frontier run without inheriting one vendor’s software stack. SF Tensor optimizes, verifies, places and operates training across the best available compute.

Plan a training runSee the training stack

Input

Training repo
Novel architecture
Scaling plan
01

Optimize

Generate fast kernels

02

Scale

1 to 10,000 GPUs

03

Recover

Resume automatically after node failure

Outcome

Frontier run
Maximum utilization
Portable training

01. Average fleet utilization

96%

Measured across a set of AI labs running SF Tensor.

02. GPUs, one workflow

1to10,000GPUs

From local research to a production pre-training fleet.

03. Failure recovery

Automatic

Resume from the last checkpoint after node failure, with no human in the loop.

The lab advantage

Frontier runs should not require frontier-lab headcount.

The scarce thing is the architecture and the people inventing it. We make compute supply fungible by retargeting workloads across vendors, compiling kernels for the actual topology and operating the fleet, so a small team can move like a large lab.

01

Bring the training repo

Keep the framework and model code your researchers already use.

02

Compile for the fleet

Search, generate and formally verify optimized kernels for each target.

03

Scale across supply

Place the job on the best available mix of hardware, providers and regions.

04

Run the experiment

Resume automatically from the last checkpoint after node failure, with no human in the loop.

What you get

The systems team behind every training run.

Use SF Tensor as the compiler, performance, distributed systems and fleet operations team that turns research code into a reliable frontier run.

Novel architectures welcome

Bring unusual layers, kernels, communication patterns and research frameworks.

Automatic kernel optimization

Generate workload-specific kernels instead of waiting months for hand-tuning.

Cross-vendor portability

Retarget training as new accelerators and lower-cost supply enter the market.

Training-native storage

Keep data and checkpoints close to workers with shared, cluster-local caching.

Resilient fleet operations

Handle placement, observability, checkpoints, node failures and run recovery.

Where it starts

Give every research direction a credible path to scale.

01

Foundation models

Pre-train any foundation model from scratch: dense, MoE, multimodal, domain-specific or something no one has named yet.

One path from ablation to full-scale training, whatever you're building.
02

Arbitrary architectures

Not a fixed menu of blessed model types. Any attention scheme, state-space model, memory system, routing or communication design, including ones that don't exist yet.

If you can define it, you can train it. No architecture left behind.
03

Alternative silicon

Use emerging accelerators when they are the best technical or economic fit.

Hardware choice without a permanent software commitment.
Also building for enterprise model teams

Your architecture is the bet. We make sure the infrastructure never is.

Bring the training repo. We optimize, place and operate the run from 1 GPU to 10,000.

Plan a training runTalk to an engineer

Train the models only you can build. One stack for enterprise post-training and frontier pre-training.

All Systems Operational

© 2026 San Francisco Tensor Company

SolutionsHomeEnterprisesAI LabsTensor CloudKernel OptimizerEmma LangTalk to an engineer
Company
Blog
Manifesto
CareersWe Are Hiring!
Engineering
SF Tensor Stamp