Joul
Production AI compute

Production AI compute, without the hyperscaler bill.

Managed graphics processing units for training, fine-tuning and inference. Describe the job; Joul matches it to the right hardware and keeps it running to the finish. At a fraction of big-cloud rates.

Tell us what you run — we will come back with a price.

You have lost a three-day run to a node that vanished at 2am. You have opened a cloud invoice that made you rethink the roadmap.That is why we built Joul.

01Where Joul sits

Reliable enough for production. Priced like it is not.

Your job runs on a pooled fleet we operate ourselves — not a single machine on somebody's shelf. You will not get a hyperscaler's enterprise service agreement from us, and you will not pay for one either.

01Runs that finish

Runs that finish.

Your work is placed on machines that suit it and watched until it is done. Read its output while it runs, stop it when you choose, and find nothing stale left behind.

02Agreed before you start

A price you agree before you start.

Your account carries an agreed ceiling: how many cards of which model, and for how long. Work inside it runs. Work beyond it is refused with a plain sentence — never queued for ever, never billed as a surprise three weeks later.

03Nothing to assemble

Nothing to assemble.

No cluster wrangling, no driver roulette. Describe the job; get a running machine.

Hyperscalers are reliable and on demand, and you pay for it. Neoclouds are reliable and affordable, and they want a long-term commitment. Marketplaces are affordable and on demand, and a node can vanish under your run. Joul is the fourth combination: turnkey, on demand, reliable and affordable, in one machine.

02Turnkey, not toolkits

Describe the job. We do the rest.

One short document — which container image, what command, how many cards of which model. The platform finds the machines, runs the work and tells you how it went.

1 /

Describe it.

Write a few lines: your image, your command, the card you want and how many. Cores and memory come with the cards, so you never calculate them.

2 /

Submit it.

One command sends the document. If your allowance covers it, it is accepted on the spot; if the fleet is busy, it waits its turn and starts when capacity frees.

3 /

It runs.

A job with several steps passes files from one step to the next on its own. A service that must stay up — a model answering requests — gets a name you choose and, when the model needs more than one machine, is split across them.

4 /

Watch it.

List everything you have running. Drill into one job part by part. Read the live output, including from a part that died and restarted. Stop anything, and everything belonging to it goes with it.

console.joul.cloud
$ cat finetune.yaml
name: finetune-support
type: workflow
steps:
  - name: train
    image: registry.example.com/team/trainer:1.4
    command: [python, train.py, --epochs, "3"]
    gpus: 4
    gpu_type: H100
# cores and memory come with the cards

One job, from a sentence to a running machine.

Marketplaces hand you parts. Joul hands you a finished machine.

03Output, not hours

Get more useful work out of every card you pay for.

An hour on a card is not an hour of useful work. Failed runs, retries, idle gaps and mismatched hardware eat a share of every invoice — and you pay for all of it, useful or not.

There is a name for the share that counts: goodput. Joul is built around it. Every job is placed on hardware that suits it — the card you named, whole machines when you ask for them — and kept doing real work, so more of what you pay for comes back as finished runs.

A typical invoiceillustrative
On Joulillustrative
Finished workFailures and retriesIdle and mismatch
The number we optimise
goodput = productive compute ÷ total compute

The share of paid compute that becomes finished work. Everything we build pushes it toward one.

04What you get

Built for teams that ship.

If you burn card-hours, Joul pays off.

Training and fine-tuningHigh-volume inferenceBatch and research workStartups watching burnTeams done gambling on spot capacity

Your image, your command.

Bring the container you already build. Nothing to rewrite.

The card you name.

Say which model your work needs and it runs only on machines that carry it. Cores and memory come with the cards.

Jobs and services alike.

Work that finishes and work that stays up, written the same way and reached under a name you choose.

Whole machines on request.

Work that must not share a machine asks for exclusive use and gets it.

Your data stays yours.

You move your own data in and out. We never fetch your inputs, never deliver your outputs, and hold none of your credentials. Your work cannot be reached by another customer's, and theirs cannot reach yours.

Sign in your way.

A password, a Google account or a GitHub account. Every request is checked at the moment it is made, so a revoked access takes effect on the very next call.

One public endpoint, one published interface, one command-line client. Your account manager can walk you through your first submission in a few minutes.

05FAQ

Straight answers.

What do I have to write?+

A short document: the container image, the command, the card model and how many. A multi-step pipeline adds a step per line. That is the whole format.

Do I need to install anything?+

One command-line client. It signs you in, submits your document, lists your runs, reads their output and stops them.

What happens when the machines are full?+

Your work waits its turn and starts when capacity frees. A busy moment is never a refusal.

What happens when I ask for more than my allowance?+

The submission is refused with a sentence saying why — which model, how many cards, or that the allowance has lapsed. Nothing is silently queued.

Can I run a service that stays up?+

Yes. Write it the same way as a job, give it a name and a port, and reach it under that name for as long as you keep it running.

Who can see my data?+

Nobody on our side. You bring the image and move the data; the platform runs the steps, holds the files passed between them and keeps the result for a retention window. It fetches nothing of yours and holds none of your credentials.

How reliable is it, really?+

More reliable than a marketplace node; a step below a hyperscaler's enterprise service agreement. For most training, fine-tuning and inference work, that is the right trade.

What is goodput?+

The share of paid compute that becomes finished work — productive compute divided by total compute. It is the number we optimise.

Get started

Tell us what you run.

Model, scale, timeline — whatever you have. Our team will come back with a price and a plan to get you running.

Enter a work email so we can reply with your price.

We read every message. No drip campaigns.

Got it — our team will be in touch shortly.Usually within one business day · hello@joul.cloud