whisper.cpp

mirror of https://github.com/ggerganov/whisper.cpp.git synced 2025-05-09 03:53:11 +00:00

Author	SHA1	Message	Date
Georgi Gerganov	6266a9f9e5	release : v1.7.2 v1.7.2	2024-11-19 18:54:22 +02:00
Stefan Sydow	d24f981fb2	sycl: fix example build (#2570 )	2024-11-18 14:57:23 +02:00
Georgi Gerganov	01d3bd7d5c	ci : use local ggml in Android build (#2567 )	2024-11-16 20:45:41 +02:00
Georgi Gerganov	bb12cd9b77	ggml : tmp workaround for whisper.cpp (skip) (#2565 )	2024-11-16 20:21:24 +02:00
Georgi Gerganov	f02b40bcb4	update : readme v1.7.2-pre	2024-11-15 16:00:10 +02:00
Georgi Gerganov	83ac2842bd	scripts : fix sync path	2024-11-15 15:24:09 +02:00
Jhen-Jie Hong	c4e95fb74d	whisper.swiftui : switch Mac dest to Mac (Designed for iPad) (#2562 )	2024-11-15 15:21:53 +02:00
Georgi Gerganov	e23721f3fb	cmake : fix ppc64 check (#0 )	2024-11-15 15:21:04 +02:00
Georgi Gerganov	c0a9f8ef85	whisper : include ggml-cpu.h (#0 )	2024-11-15 15:21:04 +02:00
Georgi Gerganov	6477b84eb6	build : fixes	2024-11-15 15:21:04 +02:00
Georgi Gerganov	24d706774d	talk-llama : sync llama.cpp	2024-11-15 15:21:04 +02:00
Georgi Gerganov	5089ab2d6a	whisper : fix build (#0 )	2024-11-15 15:21:04 +02:00
Georgi Gerganov	bdbb906817	sync : ggml	2024-11-15 15:21:04 +02:00
Alberto Cabrera Pérez	fa2ebd336e	sycl : Fixes to broken builds and test-backend-ops (llama/10257) * Fixes broken build for the SYCL CUDA backend caused by non-explicit gemm call in outprod (merged in with RWKV6 in Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration #10133) * Marks permuted MUL_MAT as unsupported to be able to run test-backend-ops * Fixes asserts in norm to fix debug builds.	2024-11-15 15:21:04 +02:00
Jeff Bolz	21b01a21b6	vulkan: Optimize contiguous copies (llama/10254) * tests: Fix memory bandwidth calculation for perf tests Add a flops calculation for flash attention. Add one GGML_OP_CPY perf test. * vulkan: Optimize contiguous copies Add a variant of the copy shader for when the tensors are contiguous. Avoid the complex addressing calculations, and do four elements per invocation to hide some other overhead. Apply similar changes to the scale shader, since scale is always contiguous. Add a "progress bar" for shader compiles.	2024-11-15 15:21:04 +02:00
Jeff Bolz	b54ce5edc5	vulkan: Throttle the number of shader compiles during the build step. (llama/10222) Fixes #9582 Spawning too many concurrent copies of glslc leads to "Failed to create pipes" errors on Linux. This change applies the same throttling we use for multithreaded pipeline creation.	2024-11-15 15:21:04 +02:00
Georgi Gerganov	26a31b78e9	metal : more precise Q*K in FA vec kernel (llama/10247)	2024-11-15 15:21:04 +02:00
Jeff Bolz	14d13c5f9f	vulkan: Fix newly added tests for permuted mul_mat and 1D im2col (llama/10226)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	5e110c2eb5	metal : reorder write loop in mul mat kernel + style (llama/10231) * metal : reorder write loop * metal : int -> short, style ggml-ci	2024-11-15 15:21:04 +02:00
Georgi Gerganov	4a9926d521	metal : fix build and some more comments (llama/10229)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	ae3c5642d0	metal : fix F32 accumulation in FA vec kernel (llama/10232)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	e287a3b627	metal : hide debug messages from normal log	2024-11-15 15:21:04 +02:00
SXX	b890243690	ggml: fix zero division in ‘dne’ calculation in CUDA COUNT_EQUAL operator when ‘ne’ is small (#10213 )	2024-11-15 15:21:04 +02:00
amritahs-ibm	b7b38f7d68	ggml : optimize llamafile cpu matrix multiplication for ppc64le (llama/10156) This change upstreams llamafile's cpu matrix multiplication kernels for ppc64le using MMA builtins for FP32 datatype. This change results in a consistent 90% improvement in input processing time, and 20% to 80% improvement in output processing time, across various batch sizes. The patch is tested with Meta-Lllama-3-8B, Mistral-7B, Llama-2-7B-chat-hf models on a IBM POWER10 machine. Signed-off-by: Amrita H S <amritahs@linux.vnet.ibm.com>	2024-11-15 15:21:04 +02:00
Georgi Gerganov	9f67aab211	metal : opt-in compile flag for BF16 (llama/10218) * metal : opt-in compile flag for BF16 ggml-ci * ci : use BF16 ggml-ci * swift : switch back to v12 * metal : has_float -> use_float ggml-ci * metal : fix BF16 check in MSL ggml-ci	2024-11-15 15:21:04 +02:00
Georgi Gerganov	8f0f785d88	metal : improve clarity (minor) (llama/10171)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	d0b8335789	metal : optimize FA kernels (llama/10171) * ggml : add ggml_flash_attn_ext_get_prec * metal : use F16 precision in FA kernels ggml-ci * metal : minor clean-up * metal : compile-guard bf16 FA kernels ggml-ci * build : remove obsolete compile flag [no ci] * metal : prevent int overflows [no ci] * cuda : disable BF16 FA ggml-ci * metal : fix BF16 requirement for FA kernels ggml-ci * make : clean-up [no ci]	2024-11-15 15:21:04 +02:00
Diego Devesa	1550be79f1	ggml : add ggml-cpu.h to the public headers (llama/10204)	2024-11-15 15:21:04 +02:00
snadampal	807f848c2f	fix q4_0_8_8 format for corrupted tokens issue (llama/10198) Co-authored-by: EC2 Default User <ec2-user@ip-172-31-62-167.us-west-2.compute.internal>	2024-11-15 15:21:04 +02:00
Zhiyuan Li	42398f13b0	Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration (llama/10133) * rwkv6: rename to wkv6 * rwkv6: support avx2 avx512 armv8 armv9 * rwkv6: update cuda file name * rwkv6: rename params * wkv on sycl * sycl: add some ops * sycl: Enhance OP support judgment * wkv6: drop armv9 and tranfer to GGML style ggml-ci * sync : ggml * update the function to use appropriate types * fix define error * Update ggml/src/ggml-cpu.c * add appropriate asserts * move element-wise functions outside * put the declaration outside the loop * rewrite to be more inline with the common pattern for distributing threads * use recommended way GGML_TENSOR_LOCALS --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Diego Devesa <slarengh@gmail.com> Co-authored-by: Plamen Minev <pacominev@gmail.com> Co-authored-by: Yuri Khrustalev <ykhrustalev@users.noreply.github.com> Co-authored-by: Meng, Hengyu <airdldl@163.com>	2024-11-15 15:21:04 +02:00
Georgi Gerganov	31c3482a4e	metal : add BF16 support (llama/8439) * ggml : add initial BF16 support ggml-ci * metal : add mul_mat_id BF16 support ggml-ci * metal : check for bfloat support on the Metal device ggml-ci * metal : better var names [no ci] * metal : do not build bfloat kernels when not supported ggml-ci * metal : try to fix BF16 support check ggml-ci * metal : this should correctly check bfloat support	2024-11-15 15:21:04 +02:00
Diego Devesa	50257af686	metal : fix from ptr buffer name (llama/10189)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	d111a0987e	ggml : adjust is_first_call init value (llama/10193) ggml-ci	2024-11-15 15:21:04 +02:00
Georgi Gerganov	915bcd2c63	metal : add quantized FA support (llama/10149) * metal : add quantized FA (vec) support ggml-ci * metal : add quantized FA (non-vec) support * metal : fix support check ggml-ci * metal : clean-up * metal : clean-up (cont) * metal : fix shared memory calc + reduce smem + comments * metal : float-correctness * metal : minor [no ci]	2024-11-15 15:21:04 +02:00
Diego Devesa	f69c8b6f1b	ggml : fix arch check in bf16_to_fp32 (llama/10164)	2024-11-15 15:21:04 +02:00
Eve	8c9044bef0	Q6_K AVX improvements (llama/10118) * q6_k instruction reordering attempt * better subtract method * should be theoretically faster small improvement with shuffle lut, likely because all loads are already done at that stage * optimize bit fiddling * handle -32 offset separately. bsums exists for a reason! * use shift * Update ggml-quants.c * have to update ci macos version to 13 as 12 doesnt work now. 13 is still x86	2024-11-15 15:21:04 +02:00
Diego Devesa	5f8e928194	ggml : fix gelu tables initialization (llama/10172)	2024-11-15 15:21:04 +02:00
Diego Devesa	25da30bd60	ggml : fix q4xx mat mul, increase ggml_aligned_malloc alignment (llama/10167)	2024-11-15 15:21:04 +02:00
snadampal	542734100e	fix build break on arm64 linux (llama/10166) This fixes the build break from the recent changes to move the CPU backend to separate files https://github.com/ggerganov/llama.cpp/pull/10144	2024-11-15 15:21:04 +02:00
Diego Devesa	b06b4c0c08	cuda : clear error after changing peer access (llama/10153)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	939d36fb4c	metal : simplify f16 and f32 dequant kernels (llama/0)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	1471e41180	metal : move dequantize templates to beginning of MSL source (llama/0)	2024-11-15 15:21:04 +02:00
leo-pony	35949192e9	CANN: adjust backend registry refactor. (llama/10158) remove buffer->iface.get_name that used in cann as it was removed in backend registry refactor PR.	2024-11-15 15:21:04 +02:00
Diego Devesa	9c817edb48	ggml : move CPU backend to a separate file (llama/10144)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	24a0feb5d9	metal : minor fixup in FA kernel (llama/10143) * metal : minor fixup in FA kernel ggml-ci * metal : use the unrolled loop variable * metal : remove unused var	2024-11-15 15:21:04 +02:00
Diego Devesa	2ab8cce7e3	llama : add simple-chat example (llama/10124) * llama : add simple-chat example --------- Co-authored-by: Xuan Son Nguyen <thichthat@gmail.com>	2024-11-15 15:21:04 +02:00
Diego Devesa	b40c255e98	llama : use smart pointers for ggml resources (llama/10117)	2024-11-15 15:21:04 +02:00
Shupei Fan	ec3e16445e	vulkan : improve ggml_vk_create_buffer error handling (llama/9898)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	0665168ef3	ggml : remove ggml_scratch (llama/10121) ggml-ci	2024-11-15 15:21:04 +02:00
Zhenwei Jin	5f6b992eea	build: fix build error in Windows env with OneAPI setup (llama/10107)	2024-11-15 15:21:04 +02:00

1 2 3 4 5 ...

1923 Commits