Run large language models locally with a single command. Download and run Llama, Mistral, Gemma and more on your own machine.
Local LLM runtime with model pulling, GPU acceleration and an OpenAI-compatible API server.