ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A high-throughput and memory-efficient inference and serving engine for LLMs
Ultralytics YOLO26, YOLO11, YOLOv8 β object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
πΈπ¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
We write your reusable computer vision tools. π
π€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
Open standard for machine learning interoperability
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
SD.Next: All-in-one WebUI for AI generative image and video creation, captioning and processing
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models