Model providers

Llama

Open-weight models you can run inside your own network.

made by
Meta

Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Meta - we just build with this.

What Llama is

Llama is Meta's family of open-weight models, released under a community licence rather than a standard open-source one. The weights can be downloaded, fine-tuned, quantized, and served on your own hardware, and the smaller sizes run on a single GPU. Most inference servers and quantization formats support them first.

How we use it

Llama is what we deploy when data cannot leave the client's network: served with vLLM on their GPUs, or with Ollama on a workstation during development. It is also the cost answer for high-volume narrow tasks - classification, extraction, routing - where a tuned small model matches a frontier model on that one job. Which one ships is decided by the client's own eval set, not by a leaderboard.

Where it is the wrong choice

On long multi-step agent work with many tools, the open models we have tested still drift and mis-call tools more often than the frontier hosted ones, so we do not put them in an autonomous loop with write access. The licence also carries conditions worth a legal read before a product ships on it.

Building something on Llama?

Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherLlama is even the right call for it.

Start a project

Llama and Meta are trademarks of their respective owners, used here to say what we work with.