whisper.cpp

mirror of https://github.com/ggerganov/whisper.cpp.git synced 2024-12-24 06:46:37 +00:00

Author	SHA1	Message	Date
Jeff Bolz	1bebb1a116	vulkan: Optimize some mat-vec mul quant shaders (llama/10296) Compute two result elements per workgroup (for Q{4,5}_{0,1}). This reuses the B loads across the rows and also reuses some addressing calculations. This required manually partially unrolling the loop, since the compiler is less willing to unroll outer loops. Add bounds-checking on the last iteration of the loop. I think this was at least partly broken before. Optimize the Q4_K shader to vectorize most loads and reduce the number of bit twiddling instructions.	2024-11-20 21:00:08 +02:00
Dan Johansson	ee437cde59	ggml : optimize Q4_0 into Q4_0_X_Y repack (llama/10324)	2024-11-20 21:00:08 +02:00
Srihari-mcw	c1506d38cf	Make updates to fix issues with clang-cl builds while using AVX512 flags (llama/10314)	2024-11-20 21:00:08 +02:00
Johannes Gäßler	c9541741e6	ggml: new optimization interface (ggml/988) * ggml: new optimization interface remove test2.c, test3.c store adamw params in tensor move grads from tensor to graph * avoid segfault upon API misuse * add ggml-opt.h to public headers * remove dependence of ggml-opt.cpp on ggml-cpu.h	2024-11-20 21:00:08 +02:00
Georgi Gerganov	6a55015dc4	ggml : remove duplicated sources from the last sync (ggml/1017) * ggml : remove duplicated sources from the last sync ggml-ci * cont : remove FindSIMD.cmake [no ci]	2024-11-20 21:00:08 +02:00
slaren	7e86030d4d	ggml : fix some build issues	2024-11-20 21:00:08 +02:00
Georgi Gerganov	401fbea326	sync : leftovers (ggml/0) ggml-ci	2024-11-20 21:00:08 +02:00
Georgi Gerganov	44d1cbdfe9	cmake : restore CMakeLists.txt (llama/10256) ggml-ci	2024-11-20 21:00:08 +02:00
Eve	3216efef2e	AVX BF16 and single scale quant optimizations (llama/10212) * use 128 bit loads (i've tried 256->128 to death and its slower) * double accumulator * avx bf16 vec dot * +3% q4_0 inference * +7% tg +5% pp compared to master * slower f16c version, kep for reference * 256b version, also slow. i tried :) * revert f16 * faster with madd * split to functions * Q8_0 and IQ4_NL, 5-7% faster * fix potential overflow (performance reduced) * 16 bit add for q4_0 only * merge	2024-11-20 21:00:08 +02:00
Romain Biessy	2c0484ebf7	sycl: Use syclcompat::dp4a (llama/10267) * sycl: Use syclcompat::dp4a * Using the syclcompat version allow the compiler to optimize the operation with native function * Update news section * Update CI Windows oneAPI version to 2025.0 * Reword doc * Call syclcompat::dp4a inside dpct::dp4a This reverts commit 90cb61d692d61360b46954a1c7f780bd2e569b73.	2024-11-20 21:00:08 +02:00
Charles Xu	3298916e5e	backend cpu: add online flow for aarch64 Q4_0 GEMV/GEMM kernels (llama/9921) * backend-cpu: add online flow for aarch64 Q4_0 GEMV/GEMM kernels --------- Co-authored-by: Diego Devesa <slarengh@gmail.com>	2024-11-20 21:00:08 +02:00
Diego Devesa	746bf2596f	ggml : build backends as libraries (llama/10256) * ggml : build backends as libraries --------- Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: R0CKSTAR <xiaodong.ye@mthreads.com>	2024-11-20 21:00:08 +02:00
Georgi Gerganov	5f7e094ccb	scripts : update sync	2024-11-20 21:00:08 +02:00
Georgi Gerganov	6266a9f9e5	release : v1.7.2	2024-11-19 18:54:22 +02:00
Stefan Sydow	d24f981fb2	sycl: fix example build (#2570 )	2024-11-18 14:57:23 +02:00
Georgi Gerganov	01d3bd7d5c	ci : use local ggml in Android build (#2567 )	2024-11-16 20:45:41 +02:00
Georgi Gerganov	bb12cd9b77	ggml : tmp workaround for whisper.cpp (skip) (#2565 )	2024-11-16 20:21:24 +02:00
Georgi Gerganov	f02b40bcb4	update : readme	2024-11-15 16:00:10 +02:00
Georgi Gerganov	83ac2842bd	scripts : fix sync path	2024-11-15 15:24:09 +02:00
Jhen-Jie Hong	c4e95fb74d	whisper.swiftui : switch Mac dest to Mac (Designed for iPad) (#2562 )	2024-11-15 15:21:53 +02:00
Georgi Gerganov	e23721f3fb	cmake : fix ppc64 check (#0 )	2024-11-15 15:21:04 +02:00
Georgi Gerganov	c0a9f8ef85	whisper : include ggml-cpu.h (#0 )	2024-11-15 15:21:04 +02:00
Georgi Gerganov	6477b84eb6	build : fixes	2024-11-15 15:21:04 +02:00
Georgi Gerganov	24d706774d	talk-llama : sync llama.cpp	2024-11-15 15:21:04 +02:00
Georgi Gerganov	5089ab2d6a	whisper : fix build (#0 )	2024-11-15 15:21:04 +02:00
Georgi Gerganov	bdbb906817	sync : ggml	2024-11-15 15:21:04 +02:00
Alberto Cabrera Pérez	fa2ebd336e	sycl : Fixes to broken builds and test-backend-ops (llama/10257) * Fixes broken build for the SYCL CUDA backend caused by non-explicit gemm call in outprod (merged in with RWKV6 in Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration #10133) * Marks permuted MUL_MAT as unsupported to be able to run test-backend-ops * Fixes asserts in norm to fix debug builds.	2024-11-15 15:21:04 +02:00
Jeff Bolz	21b01a21b6	vulkan: Optimize contiguous copies (llama/10254) * tests: Fix memory bandwidth calculation for perf tests Add a flops calculation for flash attention. Add one GGML_OP_CPY perf test. * vulkan: Optimize contiguous copies Add a variant of the copy shader for when the tensors are contiguous. Avoid the complex addressing calculations, and do four elements per invocation to hide some other overhead. Apply similar changes to the scale shader, since scale is always contiguous. Add a "progress bar" for shader compiles.	2024-11-15 15:21:04 +02:00
Jeff Bolz	b54ce5edc5	vulkan: Throttle the number of shader compiles during the build step. (llama/10222) Fixes #9582 Spawning too many concurrent copies of glslc leads to "Failed to create pipes" errors on Linux. This change applies the same throttling we use for multithreaded pipeline creation.	2024-11-15 15:21:04 +02:00
Georgi Gerganov	26a31b78e9	metal : more precise Q*K in FA vec kernel (llama/10247)	2024-11-15 15:21:04 +02:00
Jeff Bolz	14d13c5f9f	vulkan: Fix newly added tests for permuted mul_mat and 1D im2col (llama/10226)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	5e110c2eb5	metal : reorder write loop in mul mat kernel + style (llama/10231) * metal : reorder write loop * metal : int -> short, style ggml-ci	2024-11-15 15:21:04 +02:00
Georgi Gerganov	4a9926d521	metal : fix build and some more comments (llama/10229)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	ae3c5642d0	metal : fix F32 accumulation in FA vec kernel (llama/10232)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	e287a3b627	metal : hide debug messages from normal log	2024-11-15 15:21:04 +02:00
SXX	b890243690	ggml: fix zero division in ‘dne’ calculation in CUDA COUNT_EQUAL operator when ‘ne’ is small (#10213 )	2024-11-15 15:21:04 +02:00
amritahs-ibm	b7b38f7d68	ggml : optimize llamafile cpu matrix multiplication for ppc64le (llama/10156) This change upstreams llamafile's cpu matrix multiplication kernels for ppc64le using MMA builtins for FP32 datatype. This change results in a consistent 90% improvement in input processing time, and 20% to 80% improvement in output processing time, across various batch sizes. The patch is tested with Meta-Lllama-3-8B, Mistral-7B, Llama-2-7B-chat-hf models on a IBM POWER10 machine. Signed-off-by: Amrita H S <amritahs@linux.vnet.ibm.com>	2024-11-15 15:21:04 +02:00
Georgi Gerganov	9f67aab211	metal : opt-in compile flag for BF16 (llama/10218) * metal : opt-in compile flag for BF16 ggml-ci * ci : use BF16 ggml-ci * swift : switch back to v12 * metal : has_float -> use_float ggml-ci * metal : fix BF16 check in MSL ggml-ci	2024-11-15 15:21:04 +02:00
Georgi Gerganov	8f0f785d88	metal : improve clarity (minor) (llama/10171)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	d0b8335789	metal : optimize FA kernels (llama/10171) * ggml : add ggml_flash_attn_ext_get_prec * metal : use F16 precision in FA kernels ggml-ci * metal : minor clean-up * metal : compile-guard bf16 FA kernels ggml-ci * build : remove obsolete compile flag [no ci] * metal : prevent int overflows [no ci] * cuda : disable BF16 FA ggml-ci * metal : fix BF16 requirement for FA kernels ggml-ci * make : clean-up [no ci]	2024-11-15 15:21:04 +02:00
Diego Devesa	1550be79f1	ggml : add ggml-cpu.h to the public headers (llama/10204)	2024-11-15 15:21:04 +02:00
snadampal	807f848c2f	fix q4_0_8_8 format for corrupted tokens issue (llama/10198) Co-authored-by: EC2 Default User <ec2-user@ip-172-31-62-167.us-west-2.compute.internal>	2024-11-15 15:21:04 +02:00
Zhiyuan Li	42398f13b0	Optimize RWKV6 Operator Naming and Implement Multi-core CPU/ SYCL Acceleration (llama/10133) * rwkv6: rename to wkv6 * rwkv6: support avx2 avx512 armv8 armv9 * rwkv6: update cuda file name * rwkv6: rename params * wkv on sycl * sycl: add some ops * sycl: Enhance OP support judgment * wkv6: drop armv9 and tranfer to GGML style ggml-ci * sync : ggml * update the function to use appropriate types * fix define error * Update ggml/src/ggml-cpu.c * add appropriate asserts * move element-wise functions outside * put the declaration outside the loop * rewrite to be more inline with the common pattern for distributing threads * use recommended way GGML_TENSOR_LOCALS --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Diego Devesa <slarengh@gmail.com> Co-authored-by: Plamen Minev <pacominev@gmail.com> Co-authored-by: Yuri Khrustalev <ykhrustalev@users.noreply.github.com> Co-authored-by: Meng, Hengyu <airdldl@163.com>	2024-11-15 15:21:04 +02:00
Georgi Gerganov	31c3482a4e	metal : add BF16 support (llama/8439) * ggml : add initial BF16 support ggml-ci * metal : add mul_mat_id BF16 support ggml-ci * metal : check for bfloat support on the Metal device ggml-ci * metal : better var names [no ci] * metal : do not build bfloat kernels when not supported ggml-ci * metal : try to fix BF16 support check ggml-ci * metal : this should correctly check bfloat support	2024-11-15 15:21:04 +02:00
Diego Devesa	50257af686	metal : fix from ptr buffer name (llama/10189)	2024-11-15 15:21:04 +02:00
Georgi Gerganov	d111a0987e	ggml : adjust is_first_call init value (llama/10193) ggml-ci	2024-11-15 15:21:04 +02:00
Georgi Gerganov	915bcd2c63	metal : add quantized FA support (llama/10149) * metal : add quantized FA (vec) support ggml-ci * metal : add quantized FA (non-vec) support * metal : fix support check ggml-ci * metal : clean-up * metal : clean-up (cont) * metal : fix shared memory calc + reduce smem + comments * metal : float-correctness * metal : minor [no ci]	2024-11-15 15:21:04 +02:00
Diego Devesa	f69c8b6f1b	ggml : fix arch check in bf16_to_fp32 (llama/10164)	2024-11-15 15:21:04 +02:00
Eve	8c9044bef0	Q6_K AVX improvements (llama/10118) * q6_k instruction reordering attempt * better subtract method * should be theoretically faster small improvement with shuffle lut, likely because all loads are already done at that stage * optimize bit fiddling * handle -32 offset separately. bsums exists for a reason! * use shift * Update ggml-quants.c * have to update ci macos version to 13 as 12 doesnt work now. 13 is still x86	2024-11-15 15:21:04 +02:00
Diego Devesa	5f8e928194	ggml : fix gelu tables initialization (llama/10172)	2024-11-15 15:21:04 +02:00

... 2 3 4 5 6 ...

1986 Commits