Dedicated GPU capacity

GPU capacity built to your spec,
for less than any cloud.

Tell us the GPUs and the count. We deliver a dedicated cluster and run it for you.

Our team comes from Meta ByteDance Wesco Dartmouth Caltech
~30MW

GPU capacity brought online

~2,000

B300 servers delivered

17

sites for 20 enterprise customers

2X+

inference efficiency gain at Meta

Track record of Larch Labs' founders in prior roles.

Why Larch

The capacity you need, at a lower price than any cloud

Custom

The GPUs you specify, in the count and location you need.

Lower cost

A deployment fee well below any cloud's rental margin. See the math below.

Fully run

24/7 operations and performance tuning included.

The math, side by side

Typical neocloud 3-year leaseLarch
What you pay in total Hardware, site and operations, plus about 50% of the hardware cost as the cloud's return Hardware, site and operations, plus a one-time 5% of the hardware cost as initial deployment fee
Who owns the GPUs The cloud You
How it works

From request to production in three steps

  1. Send your request

    GPU model, count and timeline.

  2. Get a quote

    A fixed price and a delivery date.

  3. Go live

    We deploy, test and hand over, then keep it running.

Already own GPUs? We also operate and optimize existing clusters.

Who it's for

Teams that need serious compute

AI labs and model companies

Dedicated training and inference capacity.

Inference platforms

Lower cost per GPU-hour, so more margin per token.

GPU clouds

Build-out by people who've done it at hyperscale.

Enterprises and sovereign AI

Private capacity, delivered and run end to end.

The team

We've done this at the scale our customers are reaching for

Mark Kong

Mark Kong

Co-founder & CEO · Optimization

Mark spent more than 10 years at Meta working on AI infrastructure at production scale. He led optimization projects that more than doubled the inference GPU efficiency of a flagship AI system, and those techniques were adopted across Meta's core production fleet, saving billions of dollars in GPU spend.

At Larch Labs he leads the optimization practice: profiling customers' training and inference stacks, finding where the bottlenecks are, and lifting GPU efficiency on hardware they already own.

LinkedIn ↗
Alex S.

Alex S.

Co-founder & CTO · Deployment and Operations

Alex has 15 years of experience building infrastructure, most recently deploying GPU capacity at hyperscale. He led delivery of ~30MW and ~2,000 B300 servers across 17 sites for 20 enterprise customers, representing billions of dollars in total infrastructure value.

At Larch Labs he runs deployment and operations end to end: site selection, cluster architecture, procurement, installation and burn-in, then 24/7 operations once the cluster is live, backed by long-standing supplier relationships.

Rulan Zheng

Rulan Zheng

Co-founder & CMO · Go-to-market and BD

Rulan spent years at ByteDance building enterprise products for BytePlus, the company's B2B cloud and data business. She has also invested in early-stage AI infrastructure companies with Eastlink Capital, so she knows the market from both the builder's and the investor's side.

At Larch Labs she leads go-to-market and business development: positioning, partnerships with OEMs, colocation providers and GPU financiers, and the path from first call to signed deployment.

LinkedIn ↗
Work with us

Tell us what you need.

We'll get back to you shortly.