Open standard for machine learning interoperability
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
π€ Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A high-throughput and memory-efficient inference and serving engine for LLMs
The most powerful and modular diffusion model GUI, api and backend with a graph/nodes interface.
Stable Diffusion web UI
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
We write your reusable computer vision tools. π
ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator
Ultralytics YOLO26, YOLO11, YOLOv8 β object detection, instance segmentation, semantic segmentation, image classification, pose estimation, object tracking
π€ Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
πΈπ¬ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
SD.Next: All-in-one WebUI for AI generative image and video creation, captioning and processing
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models