Model providers

Ollama

Run open-weight models locally with one command.

made by
Ollama

Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Ollama - we just build with this.

What Ollama is

Ollama packages open-weight models together with their weights, prompt template, and parameters, and serves them on localhost over a small HTTP API that also speaks the OpenAI chat format. It handles GPU offload and quantized formats, and runs on macOS, Linux, and Windows.

How we use it

We use it for development against a local model: trying a prompt on a laptop, checking whether a small model is good enough for one step of a pipeline before committing to it, and running tests that should not spend API credits. When a proof of concept has to stay entirely offline, Ollama on a workstation is the fastest way to put something real in front of people.

Where it is the wrong choice

It is a single-node convenience runtime, not a serving stack. Throughput under concurrency is poor next to vLLM and there is no meaningful batching, autoscaling, or multi-GPU story, so anything with real traffic moves off it.

Building something on Ollama?

Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherOllama is even the right call for it.

Start a project

Ollama and Ollama are trademarks of their respective owners, used here to say what we work with.