Whisper
OpenAI · Áudio, voz e transcrição
Modelo aberto da OpenAI que transcreve e traduz fala, usado em Python ou no terminal
- 110,1 mil no GitHub
Versão em C/C++ do Whisper para transcrever áudio localmente, com GPU e quantização
por ggml.ai (Hugging Face)
Última versão: v1.9.5, de 6 de outubro de 2026
whisper.cpp é uma implementação em C/C++ do modelo de transcrição Whisper, da OpenAI, mantida pela organização ggml-org, da ggml.ai, empresa que hoje faz parte do Hugging Face. Serve a desenvolvedores que querem transcrever áudio localmente ou embutir a transcrição em outros programas.
Funciona pela linha de comando (whisper-cli) ou como biblioteca com API em C, com binários prontos para Windows e Linux. Por padrão, o whisper-cli lê só arquivos WAV de 16 bits; outros formatos exigem conversão, por exemplo com ffmpeg.
Overview
New version has been released.
Nightly build: b5454 More info: dist : releases and versioning of ggml-org projects
Changelog since v1.9.4
d1be6fde whisper : bump version to 1.9.5 (#4102) 4afec37b talk-llama : update llama.cpp to v0.6.0 3d451abe sync : ggml 47f1c7dc ggml : bump version to 0.26.0 (ggml/1652) e71b7844 CUDA: make the alloc_deps check batch independent (llama/29986) cdc87302 vulkan: fix Flash Attention shmem write out of bounds (llama/29988) 563a9f7a vulkan: revert mul_mat_id tile selection PR #29182 (llama/29936) ebb8a170 cuda: use the vector lightning indexer kernel on MUSA (llama/29990) 8c6ca4ab CUDA: Optimize accumulation in mmq for NVFP4 type (llama/29857) 82fa9df7 vulkan: fix stale prealloc_y reuse across flash attention and soft_max (llama/29591) f6915422 vulkan: sparse flash attention for quantized K/V (llama/29639) d477de3c llama : fix unexpected graph reallocation in the k-pool models (llama/29958) 95283ca7 ggml-cpu : add Q8_0 IME1 matrix kernel for SpacemiT X60 (llama/28479) 155c5c28 vulkan : Fix undeclared identifiers when -DGGML_VULKAN_RUN_TESTS=ON (llama/29912) 7091d4a2 webgpu: add MMVQ support for Q1_0/Q5_0/Q5_1/Q3_K/Q5_K/Q6_K/MXFP4 (llama/29483) 4cc62b81 cuda: tile the lightning indexer over keys and tokens for 4 heads (llama/29901) 7e576cbd metal : few-row MMA mat-mul (llama/29869) 91da4702 CUDA: use MMVF for thin f16/bf16 mul_mat at small batch size (llama/29633) 8bb85d96 CUDA: prefer whole-tile FlashAttention scheduling for efficient two-stage kernels (llama/29435) ce6552ef CUDA: refactor swizzling code (llama/29612) bfe3dc63 ggml-cpu: support BF16/FP16/FP32 K tails in tinyBLAS on x86 (llama/29806) 25453c0a cuda : move neu_padded to where it is used (llama/29940) 7914776a cuda : move blocks_per_col to where it is used (llama/29939) 6421d2c1 CUDA: fix MMQ memory fault if n_expert >> n_ubatch (llama/29941) 05d4d89a vulkan: fix rdna4 mat_vec tuning (llama/29934) 51bfbd61 webgpu: add f16 support to fill/set_rows (llama/29897) f2aa80ad ggml-openvino: update to 2026.4.1, optimize performance, expand ops, improve device listing. (llama/29852) 7f4a67a5 qwen4exp : halve the indexer score memory (llama/29825) cf5d9e30 CUDA: fuse shared experts into MMVQ (llama/29184) d7682b58 ggml-cuda : fix cpy transposed path corrupting non-contiguous dst (llama/27663) 6b704713 ggml-quants : avoid invalid rounding in qkx3 scale search (llama/29817) 0b35d1f3 ggml-cpu : fix soft_max_back wrong output when dst aliases src1 (llama/27096) d0dc4363 metal : add tensor API flash attention kernel for F16 KV (llama/29570) cca8f72b opencl: use sigmoid f16 for bf16 (llama/29787) cd1bee52 SYCL: Q8_0 DMMV ESIMD and MMVQ wide load (llama/29186) 95a05ed4 vulkan: disable large matmul tile on Samsung GPUs with 32KB shared memory (llama/28531) c4051a33 sycl: large register file for D=512 FA vec kernels (llama/29062) ae92605f sycl : do not use slow oneDNN reference matmul and fattn (llama/28985) 9c968f7e qwen4exp : optimize mask constructions (llama/29824) 4435763c ggml : add alloc_buffer_n to buffer type interface (llama/23671) 8291ab84 vulkan: add logging to pipeline compile issues (llama/29794) eaadab3d hexagon: install rebuilt HTP skels (llama/29828) 0295ef68 hexagon: add q2_k and q3_k quant type support (llama/29717) 396f68d3 CUDA: fix 2 broken Volta FA cases (llama/29803) 3a3598cc llama: refer to segment documentation [no ci] (llama/29074) 15229ac9 hexagon: shared strided DMA copy for CPY and CONCAT, any-dim CONCAT via DMA (llama/29685) aa538005 cuda : route sm70 to the Turing MMVQ nwarps table (llama/29753) 81ca4f81 metal : release temporary private transfer buffers (llama/29777) 88948bdb webgpu: add bfloat16 support for MUL_MAT/MUL_MAT_ID/GET_ROWS- #29358 (llama/29358) 10872be7 CUDA: Handle compute type for NVFP4 on cublass path (llama/29173) 3b68f901 CUDA: Make CCCL configurable + pin it to
Ainda sem avaliações
Usou o whisper.cpp? Conte para quem está escolhendo.
OpenAI · Áudio, voz e transcrição
Modelo aberto da OpenAI que transcreve e traduz fala, usado em Python ou no terminal
Chidi Williams · Áudio, voz e transcrição
Transcreve e traduz áudio offline no computador com Whisper e exporta legendas
Ficha escrita pela redação do Flipters a partir das fontes acima, revisada em 6 de outubro de 2026. Preços e versões mudam: confira no site oficial antes de comprar. Como trabalhamos.