Fast inference engine for Transformer models
A high-throughput and memory-efficient inference and serving engine for LLMs
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
Faster Whisper transcription with CTranslate2
AICI: Prompts as (Wasm) Programs