OPEN-WEIGHT MODELS COME HOME
- 4 days ago
- 1 min read

If speed was getting in the way of running open-weight models on your own hardware, it may be worth rerunning the numbers.
Model quality has improved quickly. But for some regulated workloads, another constraint remained: a sufficiently capable model could run inside the perimeter, just not fast enough to make the deployment practical.
That constraint is moving.
In May, Google released first-party multi-token prediction drafters for Gemma 4, reporting up to a 3× speedup. Then in late August, vLLM published testing across several speculative-decoding approaches and found roughly 2× gains for Gemma 4 31B across several workloads.
The mechanism is rather like giving the model a junior to draft ahead of it. The larger model verifies several proposed tokens at once rather than generating every one sequentially.
For organisations where data residency or regulatory controls rule out hosted models, some workloads that were previously impractical on owned hardware may now be worth revisiting.
In our latest essay 'When inference comes home', Dr Arsalan Shahid explains what's changed and what it could mean for regulated organisations.
Read the full article on our Substack.







