Crusoe Cloud
pricing
Flexible pricing for every workload — GPU instances, Managed Inference, and Serverless Fine-Tuning.

Built for value, scalability & speed
Cost-effective performance
Our AI-optimized hardware and lightweight virtualization cut waste and unlock performance, so you get more done with less.
Growth-aligned commitments
Avoid GPU lock-ins with tailored agreements that scale with your needs and budget.Â
Flexible consumption models
Access our full portfolio of LLMs and generative models with pay-as-you-go pricing, or utilize our GPU/CPU offerings, with spot, on-demand, or reserved pricing options.
Pricing for 
compute, inference, and fine-tuning
Pricing for compute instances and managed AI services
GPU instances pricing
Access the latest high-performance GPUs including NVIDIA GB200 NVL72 and AMD MI355X. Pay by the hour for maximum agility and unthrottled compute, or contact us to lock in guaranteed resources at our lowest rates.
CPU instances pricing
Ideal for data processing, model checkpointing and orchestrating your GPU clusters. Choose from a variety of vCPU and RAM configurations.
General-purpose
$0.04/vCPU-hr
Storage-optimized
$0.09/vCPU-hr
Storage pricing
Reliable, low-latency storage designed to handle the massive datasets and high-throughput demands of modern AI workloads.
Persistent disks
$0.08
per GiB/month
Shared disks
$0.07
per GiB/month
Container registry usage
$0.10
per GiB/month
Object Storage
$0.06
per GiB/month
Managed Kubernetes pricing
A fully managed cluster that simplifies deployment and scaling of your AI applications across GPU and CPU resources.
Cluster pricing
$0.10
per cluster hour
Severless Fine-Tuning pricing
Customize top performing models with your proprietary data.
< 16B
parameters
$0.40
Qwen3.5 2B
Qwen3.5 9B
Meta Llama 3.1 8B Instruct
Qwen3 8B
16B to 70B
parameters
$2.50
OpenAI gpt-oss 20b
Qwen3.6 35B A3B
Google gemma 4 31B it
70B to 300B
parameters
$6.00
Meta Llama 3.3 70B Instruct
OpenAI gpt-oss 120b
Qwen3 235B A22B Instruct 2507
DeepSeek V4 Flash
Serverless Inference pricing
Seamlessly integrate the industry's leading Large Language Models (LLMs) and generative models into your applications with flexible pay-as-you-go pricing.
DeepSeek
V3 0324
$0.50
$1.50
$0.25
DeepSeek
V4 Pro
$1.74
$3.48
$0.15
DeepSeek
V4 Flash
$0.14
$0.28
$0.03
Gemma 4
31B-it
$0.14
$0.40
$0.14
GLM 5.1
$1.20
$4.40
$0.25
GLM 5.2
$1.40
$4.40
$0.26
GPT-OSS
120B
$0.05
$0.20
$0.05
Kimi K2.6
$0.70
$3.50
$0.35
Llama
3.3 70B Instruct
$0.25
$0.75
$0.13
Nemotron 3 Nano
30B-A3B-FP8
$0.05
$0.20
$0.03
Nemotron 3 Nano Omni 30B A3B Reasoning
(Text, Image, Video)
$0.30
$1.83
$0.30
Nemotron 3 Nano Omni 30B A3B Reasoning
(Audio)
$0.50
$1.83
$0.50
Nemotron 3 Super
120B-A12B-FP8
$0.30
$2.40
$0.15
Qwen3
235B A22B Instruct 2507
$0.22
$0.80
$0.11
Yutori n1.5
$1.50
$5.00
$1.50
Self-Serve Deployments pricing
Spin up dedicated endpoints for open and fine-tuned models in minutes — no sales engagement required. Contact sales for monthly and volume rates.
NVIDIA H100
$5.50
NVIDIA H200
$6.00
Tailored Deployments pricing
Work directly with our team for the highest level of optimization and a dedicated, benchmarked endpoint. You bring the model and define the requirements, we handle the rest. Request benchmarking here.
Provisioned Throughput pricing
Ensure guaranteed throughput for your generative AI applications. Provisioned throughput is transacted via AI Model Units (AMUs). The longer your commitment, the lower your cost. Contact sales to learn more.
Frequently asked 
questions


On-demand pricing is our most flexible option, billed per-hour (or per-second) with no minimum commitment, and ideal for workloads where uptime, predictability, and stability are critical. Spot pricing offers significant discounts, but is better for fault-tolerant workloads that can be stopped and restarted without major disruption.
No. There are no upfront setup fees for on-demand GPU or CPU instances. Our billing is transparent; you only pay for the resources you consume for on-demand and spot consumption models.
At this time, Crusoe Cloud does not charge for network ingress or egress, either within a VPC or to/from the public internet.
Crusoe Managed Inference uses a usage-based, pay-as-you-go model, billed per 1 million tokens. Input tokens are the text your application sends to the model; Output tokens are the text the model generates in response. Cached tokens are used when the model reuses previous context or prompts, which are typically billed at a much lower rate.
Training tokens = (dataset size in tokens) × (number of epochs). You can estimate your cost before submitting a job. See our token estimation guide for details.
Serverless Inference is shared infrastructure with pay-as-you-go token pricing — ideal for experimentation and variable workloads. Self-Serve Deployments give you the option to choose the GPU and offers optimized profiles for either throughput or responsiveness, with predictable hourly billing. Use Self-Serve Deployments when you're ready to move a model into production.
Are you ready to build something amazing?
