The open-source AI voice studio. Clone, dictate, create.
A high-throughput and memory-efficient inference and serving engine for LLMs
Fast inference engine for Transformer models
Go with your own intelligence - Write Go applications that directly integrate llama.cpp for local inference using hardware acceleration on Linux, macOS, Windows, & WebAssembly.