Motor em C/C++ para rodar LLMs localmente, com CLI e servidor compatível com a OpenAI

por ggml-org (ggml.ai / Hugging Face)

Usa inteligência artificial
Nota dos leitores
Sem avaliaçõesAvaliar
No ranking
4ºIA local
Downloads pelo Flipters
0
Estrelas no GitHub
130,5 mil

Última versão: v0.6.0, de 5 de outubro de 2026

Sobre o llama.cpp

O llama.cpp é um motor de inferência em C/C++ para rodar modelos de linguagem no formato GGUF, mantido pela comunidade ggml-org com apoio da ggml.ai, empresa comprada pela Hugging Face em 2026. É a base de vários apps de IA local.

O comando llama cli conversa com modelos baixados do Hugging Face, e o llama serve sobe uma API compatível com a OpenAI e uma interface web. O site llama.app instala a CLI e oferece o app Llama para macOS e Windows 11.

Principais recursos

  • Baixa e roda modelos GGUF direto do Hugging Face
  • Servidor llama serve compatível com a API da OpenAI
  • Interface web embutida para conversar com o modelo
  • Quantização de 1,5 a 8 bits para reduzir o uso de memória
  • Backends CUDA, Metal, Vulkan, HIP e SYCL, além de CPU

Novidades

Novidades da versão v0.6.0Publicada em 5 de outubro de 2026. Notas do GitHub, em inglês.Mostrar

Overview

llama.cpp v0.6.0 introduces the new llama_batch_ext extended batch API (with llama_process) for mixed token/embedding inputs and MTP/deepstack state embeddings, adds support for the GLM-5.3-Flash (GLM5-Next) 320B hybrid model, the Clef decision model (text and vision) and MTP speculative decoding for Qwen4Exp, ships a new /v1/systemone server API for decision models (laya, julia-1, lev, openjev, kev, nimble), overhauls the Web UI with a Hugging Face Hub data layer and model download pipeline, adds a Metal tensor-API flash attention kernel for F16 KV, sparse flash attention for quantized K/V on Vulkan, and updates ggml to v0.26.0.

Highlights

  • New llama_batch_ext extended batch API with llama_process(), supporting mixed token/embedding batches and per-token "state" embeddings for MTP and deepstack models #24669
  • New models: GLM-5.3-Flash (GLM5-Next), a 320B text+vision hybrid model #27773, and the Clef decision model, fully supported with both text and vision #29831 #29969
  • Qwen4Exp: high-quality support is now available, with MTP speculative decoding (~1.5x decode speedup on DGX Spark) and various correctness fixes #29761 #29751
  • llama and server: new /v1/systemone API supporting five decision models - laya, julia-1, lev, openjev (+vision), kev #29818
  • Metal: new tensor API flash attention kernel for F16 KV #29570
  • Metal: new few-row MMA mat-mul kernels for speculative and batched decoding, up to ~3x faster mat-mul on Apple GPUs #29869
  • New llama_prefetch_rows() using MADVISE-based prefetching of PLE tensors in Qwen4Exp and Gemma4 #29599

API changes

  • include/llama.h: new llama_batch_ext batch API with llama_embd, llama_process() and llama_process_type #24669, new llama_get_causal_attn() #28876, session formats bumped to LLAMA_SESSION_VERSION 11 and LLAMA_STATE_SEQ_VERSION 4
  • include/llama-cpp.h: added llama_batch_ext_ptr and deleter for the new extended batch API #24669
  • tools/mtmd/mtmd.h: mtmd_get_memory_usage() now returns an mtmd_memory_usage struct with image_max_tokens and use_non_causal #29773
  • tools/server: new /v1/systemone endpoint for decision models #29818 and /v1/embeddings now accepts typed vision/audio/video content #29556

New models

  • GLM-5.3-Flash (GLM5-Next): 320B KDA/DSA hybrid text+vision model with mHC and MoE #27773
  • Clef decision model, fully supported with both text and vision #29831 #29969
  • Ling 3.0 VL, folded into the BailingMoeV3 architecture #29151
  • Nimble decision model #29844
  • Registered Lfm2BidirectionalForMaskedLM for LFM2.5-Encoder-230M/350M #29862
  • Added classifier_pooling support for rerankers #29627

Core changes

  • Migrated examples, speculative decoding, mtmd and server to the new `ll
Ver a versão no GitHub

Avaliações de quem usa

Ainda sem avaliações

Usou o llama.cpp? Conte para quem está escolhendo.

Avalie o llama.cpp

Sua nota (obrigatória)

De 20 a 3.000 caracteres.0/3.000

Alternativas ao llama.cpp

  • Ollama

    Ollama Inc. · Rodar IA no seu computador

    Baixa e roda modelos de linguagem abertos no computador, por app, terminal ou API local

    • 182,4 mil no GitHub
  • LocalAI

    Ettore Di Giacinto e equipe LocalAI · Rodar IA no seu computador

    Servidor de IA local com API compatível com a OpenAI para texto, voz, imagem e vídeo

    • 49,4 mil no GitHub
  • LM Studio

    Element Labs, Inc. · Rodar IA no seu computador

    App de desktop para baixar e rodar modelos de linguagem no PC, com chat e API local

Ver todos em Rodar IA no seu computador

Fontes

Ficha escrita pela redação do Flipters a partir das fontes acima, revisada em 6 de outubro de 2026. Preços e versões mudam: confira no site oficial antes de comprar. Como trabalhamos.