Model providers

vLLM

Serving stack for open-weight models on your own GPUs.

made by
vLLM project

Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of vLLM project - we just build with this.

What vLLM is

vLLM is an inference server for open-weight models. Continuous batching and paged attention keep GPU memory and throughput high under concurrency, and it exposes an OpenAI-compatible HTTP API so existing clients work unchanged. It supports tensor parallelism across GPUs, quantized weights, and constrained output.

How we use it

When inference has to run inside the client's own network, this is what we deploy: a container per model behind their load balancer, with the OpenAI-compatible endpoint so application code does not care where the model lives. Batch and context settings are sized against their actual traffic shape rather than a published benchmark, and a hosted provider stays configured as a fallback path. Its Prometheus metrics feed the same dashboards as everything else.

Where it is the wrong choice

It assumes GPUs you have already paid for and someone to keep drivers and model versions current; below steady utilization a hosted API is cheaper and much less work. Loading a large model takes minutes, so scale-to-zero is not on the table.

Building something on vLLM?

Send the problem rather than a job spec. You get an answer on scope, on fit, and on whethervLLM is even the right call for it.

Start a project

vLLM and vLLM project are trademarks of their respective owners, used here to say what we work with.