Fractional dedicated compute
A fraction of the GPU.
Inside your dedicated environment.
Give smaller workloads an isolated slice of a real GPU, sized to what the work needs, on hardware reserved for your organization or a group you choose. Not a container dropped onto a machine shared with strangers.
The allocation problem
A workload should not tie up capacity it cannot use.
Across the clusters anyone has measured properly, machines are busy somewhere between a few percent and the low teens. Even clusters that count as well run commonly sit below a third. Meanwhile the same organizations wait in line for capacity and sign purchase orders for more.
A notebook, a smaller model, or a light inference service may need part of a GPU. Give each one a whole device and capacity sits idle while other users wait. Start with what today's project needs, not a year-long reservation sized for a roadmap that has not happened yet.
How it works
Fractional allocation. Dedicated hardware.
Fractional describes what a workload receives.
An isolated slice of a real GPU, sized to what the work actually needs, so several lighter workloads share the device without getting in each other's way.
Dedicated describes who is allowed on the hardware.
The machine is reserved for your organization or a group you choose, and your administrator picks who those people are.
Because the slice belongs to one customer, speed stops being a lottery. We will not put your work next to a stranger's work.
The operating sequence
Allocate, share, and reclaim.
Define what the workload needs.
GPU memory, GPU count, cores, RAM, and storage, set by the person running it or by policy.
Set the limits, enforced at launch.
The administrator caps what any person or group can use, and the system holds them to it the moment work starts, not on a bill.
Share by fair share when the machines fill up.
Light users run in slices of one card. When the queue is full, whoever has used the least goes next, not whoever asked first.
Warn a quiet session, then take the machine back.
Capacity returns to whoever is waiting, not to whoever forgot.
Fair use inside the system
Fair use has to live inside the system.
The tools most companies use assume a single tenant who owns the machines, so they have no concept of fairness between the people competing for them. Inside a Grid, an administrator sets the rules, and the system holds people to them the moment work starts, instead of letting them find out on a bill.
Limits per person or group
Memory, number of GPUs, processor cores, RAM, storage, and a monthly budget of GPU hours. Set before launch, enforced at launch.
One GPU, several slices
A single card splits into separate slices, so several light users share it without getting in each other’s way.
Fair-share queuing
When the machines fill up, work waits in line by fair share. Whoever has used the least goes next rather than whoever asked first.
Idle sessions warned, then reclaimed
The system warns a session that has gone quiet, then takes the machine back.
Production kept apart from training
The machines serving live traffic stay separate from the ones running training, so a production service and an experiment never compete for the same memory.
Notice what the administrator controls. Not only how much, but who. Their people share their machines, and the administrator picks who those people are.
Three questions
What buyers ask about fractional.
Put the right amount of GPU behind each workload.
Tell us what you want to run and how your team will share it. We will help you plan suitable Kinesis capacity and allocations.