Storage
发布时间:2026-10-07 | 浏览:1
Store models, datasets, and artifacts with simple per-TB pricing. Built-in CDN, Xet deduplication, and no git overhead.
Trusted by more than 10,000 AI teams
Storage built for AI teams
Store models, datasets, and artifacts with simple per-TB pricing. Xet deduplication. Included CDN. No git overhead.
Per-TB pricing with built-in CDN and deduplication speedups.
Per-TB pricing with built-in CDN and deduplication speedups.
No Git constraints: commit-free sync and fast object updates.
No Git constraints: commit-free sync and fast object updates.
Designed for ML workflows: datasets, checkpoints, model artifacts.
Designed for ML workflows: datasets, checkpoints, model artifacts.
Next-gen large-scale storage for AI
Xet uses content-defined chunking to break files into byte-level chunks and deduplicates across your entire bucket. When you retrain a model and only 5% of weights change, only that 5% is re-uploaded.
Raw + processed dataset: stored once, billed once*
Raw + processed dataset: stored once, billed once*
4x less data per upload, verified with real-world workloads
4x less data per upload, verified with real-world workloads
*Requires Enterprise or Enterprise Plus plan
Traditional S3 Upload
8 / 8 chunks uploaded
XET Deduplicated Upload
1 / 8 chunks uploaded
Gray = already stored · Purple = only the changed chunk
Transparent, volume-based pricing
Simple per-TB pricing that scales with usage. Egress and CDN are included at no extra cost.
Assemble training data at any scale
Pour raw data from every source into a single bucket: crawls, annotations, synthetic outputs, partner datasets. No git overhead, no commit queues, no file-count limits. When training begins, your data is already there, streamed to GPUs via the included CDN.
Immediate availability on upload, no queued commits
Immediate availability on upload, no queued commits
Batch API processes thousands of files in a single call
Batch API processes thousands of files in a single call
Raw + processed datasets with dedup = no double billing*
Raw + processed datasets with dedup = no double billing*
*Requires Enterprise or Enterprise Plus plan
crawl-2026-jan/
48 TB · 2.1M files
annotations-v3/
12 TB · 890K files
synthetic-pairs/
6 TB · 340K files
Xet dedup: 66 TB stored → billed for 41 TB*
Your data, independent of your compute
Your data lives in one neutral home, not inside a single cloud. Point training and inference at whichever provider has capacity or the best price, and switch without re-uploading petabytes or paying egress to leave.
Train on AWS, GCP, Nebius, or your own cluster from one bucket
Train on AWS, GCP, Nebius, or your own cluster from one bucket
Never locked to one provider's prices or capacity
Never locked to one provider's prices or capacity
Pre-warmed CDN keeps data next to your GPUs, wherever they run
Pre-warmed CDN keeps data next to your GPUs, wherever they run
hf://buckets/acme · 240 TB
Built-in CDN for blazing fast access
Every bucket includes a CDN. Warm localized cache close to where you compute for ultra fast streaming and downloads. Egress is included up to a generous 8:1 ratio of your total storage.
Pre-warm cache in any cloud region you need
Pre-warm cache in any cloud region you need
Our CDN is deployed inside GCP and AWS networks
Our CDN is deployed inside GCP and AWS networks
Egress included up to 8:1 your storage
Egress included up to 8:1 your storage
Give your coding agents persistent storage
Coding agents run in ephemeral environments, but their outputs shouldn't vanish. Checkpoints, benchmark results, generated datasets: one hf sync command in your agent's bash tool is all it takes.
Pre-warmed CDN and no git overhead for fast reads and writes
Pre-warmed CDN and no git overhead for fast reads and writes
Persist artifacts across ephemeral CI runs and terminal sessions
Persist artifacts across ephemeral CI runs and terminal sessions
Install the official HF CLI skill and your agent knows every command
Install the official HF CLI skill and your agent knows every command
Enterprise-grade security at every layer
AES-256 Encryption
End-to-end encryption at rest and in transit
Full visibility into every access event
Enterprise SSO with role-based access control
US & EU Regions
Choose where your data lives
Get started with HF Storage Buckets HF Storage Buckets
Start with buckets, sync your AI data, and unlock object storage built for ML workflows.