Datasets next to the GPUs.
Fine-tuning corpora and batch inputs are large and read many times. Keep them in the same cloud and region as the jobs that read them.
datasetsGCGCS · us-central1Documents and audio come in. Transcripts, summaries, images and embeddings go out. Uplint gives every one of them a stable identity, a home next to the model that needs it, and a record of what it came from — so your pipeline is about the model, not the storage.
Outputs are uploaded like any file, with the input, the job and the model in their metadata. A month later, “which document produced this summary?” is one lookup — not an archaeology project.
file_8Kx92mcontract-analyser-v3job_4471 · 41 s{id: "file_Q1m8ra",name: "summary.md",size: "6 KB",storage: "artifacts", // R2 · edgemetadata: {derivedFrom: "file_8Kx92m",job: "job_4471",model: "contract-analyser-v3",tenant: "acme"}}derivedFromGPU jobs read large inputs and write large outputs. One routing policy per target keeps them in the same cloud and region as the model — and puts generated media at the edge, where users are.
Fine-tuning corpora and batch inputs are large and read many times. Keep them in the same cloud and region as the jobs that read them.
datasetsGCGCS · us-central1Images, audio and video your models produce are served to users, often many times. Land them where delivery is cheap and close.
generationsR2R2 · globalContracts and records your customers upload carry residency rules. Route them by tenant; the pipeline fetches by ID and never learns the bucket.
documentsAZper tenant · policyAI products create far more files than they keep. Promote the ones users saved, cool the rest, delete on schedule — all against the same IDs, so nothing your app references ever breaks.
Every output is uploaded with lineage. Cheap to create, cheap to keep for now.
POST /v1/files · storage: "generations"The app hands out short-lived signed URLs. Nothing is public; nothing is copied.
POST /v1/files/:id/urlSaved generations get promoted to a durable target. Same ID, so the link in the user’s history still works.
POST /v1/files/:id/move → "library"Everything else moves to cold storage after 30 days and is deleted after 90 — with an event on the record.
DELETE /v1/files/:idFetch the input by ID, run the model, upload the result with its lineage. The worker is the same whether it runs on your GPUs, a managed endpoint, or a laptop.
const { url } = await uplint.files.url(job.inputId, { expires: 600 });const result = await model.analyse(await fetch(url));const artifact = await uplint.files.upload({file: result.summary,storage: "artifacts",metadata: { derivedFrom: job.inputId, job: job.id, model: "contract-analyser-v3" }});// artifact.id goes back to the app — it never sees a bucket, a region or a providerNo. Uplint stores, identifies, routes and serves files. Your workers do the extraction, transcription and generation — Uplint is the layer they read inputs from and write outputs to, so none of them need to know a bucket.
Put the input’s file ID in the output’s metadata when you upload it — derivedFrom, plus the job and model if you want them. Any artifact can then be traced back with one GET, and any input’s outputs found by querying on it.
Yes. Give datasets their own storage target and point its policy at the cloud and region where the jobs run. Uploads land there; the workers fetch by ID from the same place.
Treat them as a lifecycle, not a leak: serve them through expiring URLs, promote the ones users save with a move, and delete the rest on a schedule. File IDs stay valid until you delete them, so your app’s history never points at a missing object.
Yes — route by tenant, or connect the customer’s bucket as a storage target. The pipeline still fetches by ID and uploads by target, so nothing in the worker changes.
Connect the buckets your jobs already use, point a target at each one, and give every input and output an ID your pipeline can trust.