Llama
Open-weight models you can run inside your own network.
- made by
- Meta
- source
- Official Llama site
Magic Ship is one shop in Vancouver, BC, working remotely with clients worldwide. We are not a partner, reseller, or certified vendor of Meta - we just build with this.
What Llama is
Llama is Meta's family of open-weight models, released under a community licence rather than a standard open-source one. The weights can be downloaded, fine-tuned, quantized, and served on your own hardware, and the smaller sizes run on a single GPU. Most inference servers and quantization formats support them first.
How we use it
Llama is what we deploy when data cannot leave the client's network: served with vLLM on their GPUs, or with Ollama on a workstation during development. It is also the cost answer for high-volume narrow tasks - classification, extraction, routing - where a tuned small model matches a frontier model on that one job. Which one ships is decided by the client's own eval set, not by a leaderboard.
Where it is the wrong choice
On long multi-step agent work with many tools, the open models we have tested still drift and mis-call tools more often than the frontier hosted ones, so we do not put them in an autonomous loop with write access. The licence also carries conditions worth a legal read before a product ships on it.
Service lines it turns up in
Related tools
More in Model providers
Other tools in the same service lines
Building something on Llama?
Send the problem rather than a job spec. You get an answer on scope, on fit, and on whetherLlama is even the right call for it.
Start a projectLlama and Meta are trademarks of their respective owners, used here to say what we work with.