Install this package:
emerge -a sci-misc/llama-cpp
| Version | EAPI | Keywords | Slot |
|---|---|---|---|
| 9999 | 8 | 0 | |
| 0.6.0_p11435 | 8 | 0 | |
| 0.6.0 | 8 | ~amd64 ~arm64 | 0 |
| 0.5.0_p11398 | 8 | 0 | |
| 0.5.0_p11371 | 8 | 0 | |
| 0.5.0 | 8 | ~amd64 ~arm64 | 0 |
| 0.4.1 | 8 | ~amd64 ~arm64 | 0 |
<pkgmetadata>
<maintainer type="person">
<email>iohann.s.titov@gmail.com</email>
<name>Ivan S. Titov</name>
</maintainer>
<longdescription>
llama.cpp runs LLaMA-family models in C/C++ on CPU and on several GPU
backends. This ebuild was forked from the ::guru overlay at build
b9090 and has diverged since.
Versions follow ::guru's naming. A release keeps upstream's own
X.Y.Z and is keyworded. A snapshot of upstream build bN is named
X.Y.Z_pN after the release it follows, so it sorts above that
release and below the next one; it tracks the tip between releases
and stays unkeyworded, so it has to be requested per version through
package.accept_keywords.
</longdescription>
<use>
<flag name="blis">Build a BLIS backend</flag>
<flag name="flexiblas">Build a FlexiBLAS backend</flag>
<flag name="rocm">Build a HIP (ROCm) backend</flag>
<flag name="openblas">Build an OpenBLAS backend</flag>
<flag name="opencl">Build an OpenCL backend, so far only works on Adreno and Intel GPUs</flag>
<flag name="openssl">Use openssl to support HTTPS</flag>
<flag name="sycl">Build an Intel SYCL backend (Arc GPU, Intel CPU via
oneAPI). Requires a -fsycl-capable compiler (Intel icpx or clang++
with SYCL patches) installed separately.</flag>
<flag name="webui">Embed the llama-server web UI, using the prebuilt
assets upstream publishes with each release (the live ebuild builds
or downloads them at configure time); disable for a server binary
with only the HTTP API.</flag>
</use>
<upstream>
<remote-id type="github">ggml-org/llama.cpp</remote-id>
</upstream>
</pkgmetadata>
Manage flags for this package:
euse -i <flag> -p sci-misc/llama-cpp |
euse -E <flag> -p sci-misc/llama-cpp |
euse -D <flag> -p sci-misc/llama-cpp
| Flag | Description | 9999 | 0.6.0_p11435 | 0.6.0 | 0.5.0_p11398 | 0.5.0_p11371 | 0.5.0 | 0.4.1 |
|---|---|---|---|---|---|---|---|---|
| ( | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| ) | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| amx_bf16 | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| amx_int8 | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| amx_tile | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| avx | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| avx2 | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| avx512_bf16 | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| avx512_vnni | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| avx512f | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| avx512vbmi | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| avx_vnni | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| blis | Build a BLIS backend | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| bmi2 | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| cuda | Build the CUDA (cuda_v13) llama-server GPU backend via <pkg>dev-util/nvidia-cuda-toolkit</pkg> ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| debug | Enable debug messages ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| examples | Build and install the example programs ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| f16c | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| flexiblas | Build a FlexiBLAS backend | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| fma3 | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| openblas | Build an OpenBLAS backend | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| opencl | Build an OpenCL backend, so far only works on Adreno and Intel GPUs | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| openmp | Build the host backend with OpenMP threading ⚠️ | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ |
| openssl | Use openssl to support HTTPS | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ |
| rocm | Build a HIP (ROCm) backend | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| sse4_2 | ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| sycl | Build an Intel SYCL backend (Arc GPU, Intel CPU via oneAPI). Requires a -fsycl-capable compiler (Intel icpx or clang++ with SYCL patches) installed separately. | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| vulkan | Build the Vulkan llama-server GPU backend ⚠️ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| webui | Embed the llama-server web UI, using the prebuilt assets upstream publishes with each release (the live ebuild builds or downloads them at configure time); disable for a server binary with only the HTTP API. | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ | ⊕ |
| Type | File | Size | Versions |
|---|---|---|---|
| DIST | ggml-org_models_tinyllamas_stories15M-q4_0-99dd1a73db5a37100bd4ae633f4cfce6560e1567.gguf | 19077344 bytes | 9999, 0.6.0_p11435, 0.6.0, 0.5.0_p11398, 0.5.0_p11371, 0.5.0, 0.4.1 |
| DIST | llama-b10964-ui.tar.gz | 3084524 bytes | 0.4.1 |
| DIST | llama-b11146-ui.tar.gz | 3085812 bytes | 0.5.0 |
| DIST | llama-b11371-ui.tar.gz | 3100970 bytes | 0.5.0_p11371 |
| DIST | llama-b11398-ui.tar.gz | 3099976 bytes | 0.5.0_p11398 |
| DIST | llama-b11429-ui.tar.gz | 3099802 bytes | 0.6.0 |
| DIST | llama-b11435-ui.tar.gz | 3100086 bytes | 0.6.0_p11435 |
| DIST | llama-cpp-0.4.1.tar.gz | 37422659 bytes | 0.4.1 |
| DIST | llama-cpp-0.5.0.tar.gz | 37561391 bytes | 0.5.0 |
| DIST | llama-cpp-0.5.0_p11371.tar.gz | 37802235 bytes | 0.5.0_p11371 |
| DIST | llama-cpp-0.5.0_p11398.tar.gz | 37845371 bytes | 0.5.0_p11398 |
| DIST | llama-cpp-0.6.0.tar.gz | 37874349 bytes | 0.6.0 |
| DIST | llama-cpp-0.6.0_p11435.tar.gz | 37887588 bytes | 0.6.0_p11435 |
| Type | File | Size |
|---|