whisper.cpp

mirror of https://github.com/ggerganov/whisper.cpp.git synced 2024-12-25 07:01:04 +00:00

Author	SHA1	Message	Date
Jeroen Mostert	86506b0c5c	Allow all RDNA2 archs to use sdot4 intrinsic (llama/8629) The check gating the use of `__builtin_amdgc_sdot4` specifically checks for gfx1030. This causes a severe perf regression for anything gfx103? that's not gfx1030 and not using `HSA_OVERRIDE_GFX_VERSION` (if you've built ROCm to support it). We already have a generic RDNA2 define, let's use it.	2024-08-08 22:48:46 +03:00
luoyu-intel	11182fae34	fix scratch size of softmax (llama/8642)	2024-08-08 22:48:46 +03:00
Mark Zhuang	0bc8bffe1d	ggml: fix compile error for RISC-V (llama/8623)	2024-08-08 22:48:46 +03:00
Johannes Gäßler	8c4f30497a	CUDA: MMQ code deduplication + iquant support (llama/8495) * CUDA: MMQ code deduplication + iquant support * 1 less parallel job for CI build	2024-08-08 22:48:46 +03:00
Georgi Gerganov	b1ee3a8444	gguf : handle null name during init (llama/8587)	2024-08-08 22:48:46 +03:00
slaren	be9a16fd3f	ggml : fix quant dot product with odd number of blocks (llama/8549) * ggml : fix iq4_nl dot product with odd number of blocks * ggml : fix odd blocks for ARM_NEON (llama/8556) * ggml : fix iq4_nl dot product with odd number of blocks * ggml : fix q4_1 * ggml : fix q5_0 * ggml : fix q5_1 * ggml : fix iq4_nl metal ggml-ci * ggml : fix q4_0 * ggml : fix q8_0 ggml-ci * ggml : remove special Q4_0 code for first 2 blocks * ggml : fix sumf redefinition --------- Co-authored-by: slaren <slarengh@gmail.com> --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2024-08-08 22:48:46 +03:00
Clint Herron	f4d9a95b0f	ggml : add friendlier error message to fopen errors (llama/8575) * Add additional error information when model files fail to load. * Adding additional error information to most instances of fopen.	2024-08-08 22:48:46 +03:00
Johannes Gäßler	a8ab3abe09	CUDA: fix partial offloading for ne0 % 256 != 0 (llama/8572)	2024-08-08 22:48:46 +03:00
65a	fb6a835938	cmake : install all ggml public headers (llama/8480) Co-authored-by: 65a <65a@65a.invalid>	2024-08-08 22:48:46 +03:00
hipudding	8923bb4292	Add Ascend NPU backend (llama/6035) * [CANN] Add Ascend NPU backend Ascend is a full-stack AI computing infrastructure for industry applications and services based on Huawei Ascend processors and software. CANN (Compute Architecture of Neural Networks), developped by Huawei, is a heterogeneous computing architecture for AI. Co-authored-by: wangshuai09 <391746016@qq.com> * delete trailing whitespaces * Modify the code based on review comment * Rename LLAMA_CANN to GGML_CANN * Make ggml-common.h private * add ggml_cann prefix for acl funcs * Add logging for CANN backend * Delete Trailing whitespace --------- Co-authored-by: wangshuai09 <391746016@qq.com>	2024-08-08 22:48:46 +03:00
Johannes Gäßler	fcba6aa352	make/cmake: add missing force MMQ/cuBLAS for HIP (llama/8515)	2024-08-08 22:48:46 +03:00
Xuan Son Nguyen	8807fe608b	Refactor lora adapter support (llama/8332) * lora: load to devide buft * add patch tensor function * correct tensor patch * llama_lora_adapter_apply * correct ggml_backend_tensor_copy * add llm_build_mm * fix auto merge * update based on review comments * add convert script * no more transpose A * add f16 convert * add metadata check * add sanity check * fix ftype * add requirements * fix requirements * fix outfile * conversion: only allow selected models * fix types * cuda : do not use dmmv if the tensor does not have enough cols * llama : lora fixes * do not disable mmap with lora Co-authored-by: slaren <slarengh@gmail.com> * llm_build_lora_mm_id * convert_lora : MoE LoRA conversion support * convert_lora : prefer safetensors, similarly to convert_hf * convert_hf : simplify modify_tensors for InternLM2 * convert_lora : lazy conversion * llama : load and use alpha from LoRA adapters * llama : use llm_build_lora_mm in most model graphs * auto scale * Revert "auto scale" This reverts commit 42415a4874e0f963e4aca6796ea5dfb97cd17464. * remove redundant params * Apply suggestions from code review Co-authored-by: slaren <slarengh@gmail.com> * change kv metadata * move add_type to __init__ * convert_hf : move add_type to main() * convert_lora : use the GGUFWriter from Model instead of overwriting it --------- Co-authored-by: slaren <slarengh@gmail.com> Co-authored-by: Francis Couture-Harpin <git@compilade.net>	2024-08-08 22:48:46 +03:00
Meng, Hengyu	3e94c7a81d	add concat through dim 1/2 (llama/8483) * add concat through dim 1/2	2024-08-08 22:48:46 +03:00
0cc4m	77af3254e1	Vulkan MMQ Fix (llama/8479) * Fix incoherence by adding missing LOAD_VEC_A parameter * Fix Vulkan op result checker build error	2024-08-08 22:48:46 +03:00
bandoti	d4b3cffec4	vulkan : cmake integration (llama/8119) * Add Vulkan to CMake pkg * Add Sycl to CMake pkg * Add OpenMP to CMake pkg * Split generated shader file into separate translation unit * Add CMake target for Vulkan shaders * Update README.md * Add make target for Vulkan shaders * Use pkg-config to locate vulkan library * Add vulkan SDK dep to ubuntu-22-cmake-vulkan workflow * Clean up tabs * Move sudo to apt-key invocation * Forward GGML_EXTRA_LIBS to CMake config pkg * Update vulkan obj file paths * Add shaderc to nix pkg * Add python3 to Vulkan nix build * Link against ggml in cmake pkg * Remove Python dependency from Vulkan build * code review changes * Remove trailing newline * Add cflags from pkg-config to fix w64devkit build * Update README.md * Remove trailing whitespace * Update README.md * Remove trailing whitespace * Fix doc heading * Make glslc required Vulkan component * remove clblast from nix pkg	2024-08-08 22:48:46 +03:00
Georgi Gerganov	b852a4c5ca	metal : template-ify some of the kernels (llama/8447) ggml-ci	2024-08-08 22:48:46 +03:00
Georgi Gerganov	2157abaab4	ggml : minor naming changes (llama/8433) * ggml : minor naming changes ggml-ci * ggml : use PRId64 [no ci] * ggml : revert FA K/Q names	2024-08-08 22:48:46 +03:00
Chen Xi	68d609a12c	fix the mul_mat_id ut issues (llama/8427) * fix part of mul_mat_id * skip the bfloat 16 sycl ut Signed-off-by: Chen Xi <xi2chen@intel.com> --------- Signed-off-by: Chen Xi <xi2chen@intel.com> Co-authored-by: Meng, Hengyu <hengyu.meng@intel.com> Co-authored-by: Chen Xi <xi2chen@intel.com>	2024-08-08 22:48:46 +03:00
Nicholai Tukanov	5a8ae474f0	ggml : add NVPL BLAS support (ggml/8329) (llama/8425) * ggml : add NVPL BLAS support * ggml : replace `<BLASLIB>_ENABLE_CBLAS` with `GGML_BLAS_USE_<BLASLIB>` --------- Co-authored-by: ntukanov <ntukanov@nvidia.com>	2024-08-08 22:48:46 +03:00
Daniel Bevenius	84493d7f3e	cuda : suppress 'noreturn' warn in no_device_code (llama/8414) * cuda : suppress 'noreturn' warn in no_device_code This commit adds a while(true) loop to the no_device_code function in common.cuh. This is done to suppress the warning: ```console /src/ggml-cuda/template-instances/../common.cuh:346:1: warning: function declared 'noreturn' should not return [-Winvalid-noreturn] 346 \| } \| ^ ``` The motivation for this is to reduce the number of warnings when compilng with GGML_HIPBLAS=ON. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> * squash! cuda : suppress 'noreturn' warn in no_device_code Update __trap macro instead of using a while loop to suppress the warning. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> --------- Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2024-08-08 22:48:46 +03:00
Johannes Gäßler	15d71189e9	CUDA: optimize and refactor MMQ (llama/8416) * CUDA: optimize and refactor MMQ * explicit q8_1 memory layouts, add documentation	2024-08-08 22:48:46 +03:00
AidanBeltonS	37e962580f	Use multi_ptr to clean up deprecated warnings (llama/8256)	2024-08-08 22:48:46 +03:00
Georgi Gerganov	db0ea7a2f2	ggml : move sgemm sources to llamafile subfolder (llama/8394) ggml-ci	2024-08-08 22:48:46 +03:00
Dibakar Gope	5498b0e6c0	ggml : add AArch64 optimized GEMV and GEMM Q4 kernels (llama/5780) * Arm AArch64: optimized GEMV and GEMM kernels for q4_0_q8_0, and q8_0_q8_0 quantization * Arm AArch64: add optimized GEMV and GEMM asm kernels for q4_0_q8_0 quantization and refactor code to address llama.cpp pr#5780 suggestions * Arm AArch64: add optimized GEMV and GEMM asm kernels for q4_0_q8_0 quantization and refactor code to address llama.cpp pr#5780 suggestions * Arm AArch64: add optimized GEMV and GEMM asm kernels for q4_0_q8_0 quantization and refactor code to address llama.cpp pr#5780 suggestions * Arm AArch64: add optimized GEMV and GEMM asm kernels for q4_0_q8_0 quantization and refactor code to address llama.cpp pr#5780 suggestions * Arm AArch64: add copyright claim only to ggml-aarch64.cpp and ggml-aarch64.h files * Arm AArch64: minor code refactoring for rebase * Arm AArch64: minor code refactoring for resolving a build issue with cmake * Arm AArch64: minor code refactoring to split the Q4_0_AARC64 type into three separate types: Q4_0_4_4, Q4_0_4_8, and Q4_0_8_8 * Arm AArch64: minor code change for resolving a build issue with server-windows * retrigger checks * Arm AArch64: minor code changes for rebase * Arm AArch64: minor changes to skip the pr#7433 vec_dot code for arm cpus with SVE VL not equal to 256 bits * Arm AArch64: remove stale LLAMA_QKK_64 from CMakeLists.txt and delete build.zig * Arm AArch64: add reference scalar gemm and gemv, and avoid dynamic memory allocations during quantization for Q4_0_4_4, Q4_0_4_8, and Q4_0_8_8 * Arm AArch64: add multithreaded quantization support for the new types: Q4_0_4_4, Q4_0_4_8, and Q4_0_8_8 * Arm AArch64: minor code refactoring * Arm AArch64: simplify logic for calling gemm and gemv functions in ggml_compute_forward_mul_mat * Arm AArch64: minimize changes in ggml_compute_forward_mul_mat * Arm AArch64: minor code refactoring, and add reference scalar code to quantize routines for new quant types * Arm AArch64: minor code refactoring * Arm AArch64: minor code refactoring * Arm AArch64: minor code refactoring * rebase on the latest master commit 3fd62a6 and adapt to the new directory structure * Arm AArch64: remove a redundant comment * Arm AArch64: add pragma in ggml-aarch64.c to turn -Woverlength-strings warning off * Arm AArch64: use __aarch64__ check to guard 64-bit neon kernels * Arm AArch64: update docs/build.md README to include compile time flags for buiilding the Q4_0_4_4 quant type	2024-08-08 22:48:46 +03:00
Alberto Cabrera Pérez	2af4a52c39	sycl : Reenabled mmvq path for the SYCL Nvidia Backend (llama/8372) * SYCL : Reenabled mmvq path for the SYCL Nvidia Backend * Reduced verbosity of comment	2024-08-08 22:48:46 +03:00
Alberto Cabrera Pérez	eee2fe882e	sycl : fix powf call in device code (llama/8368)	2024-08-08 22:48:46 +03:00
Mahesh Madhav	0d1a11e5e2	ggml : loop tiling optimizations for scalar path (ggml/898) Apply a loop tiling technique to the generic path, which provides performance upside for ISAs with enough registers to take advantage of it. Also helps the compiler optimize this path.	2024-08-08 22:48:46 +03:00
Ivan Filipov	b2ead7d6f4	ggml: add support for float16 input tensors in pooling operations (ggml/895) * Add support for float16 tensors in 1d pooling operations * Add support for float16 input tensors in 2d pooling operations * code cleanup remove unnecessary casting during srow ptr initialization --------- Co-authored-by: vanaka11 <vanaka1189@gmail.com>	2024-08-08 22:48:46 +03:00
Tony Wasserka	8da6fd4dff	vulkan : initialize vk_buffer_struct members to VK_NULL_HANDLE (ggml/893) This prevents invalid frees when destroying a partially initialized vk_buffer_struct. For example, this could happen in ggml_vk_create_buffer when running out of device memory. Co-authored-by: Tony Wasserka <neobrain@users.noreply.github.com>	2024-08-08 22:48:46 +03:00
Borislav Stanimirov	ab8ec9e940	cmake : only enable GGML_NATIVE and x86 flags if not crosscompiling (ggml/885)	2024-08-08 22:48:46 +03:00
Georgi Gerganov	701265bf38	scripts : sync new files (#0 )	2024-08-08 22:48:46 +03:00
Daven Sanassy	fe36c90971	cmake : fix compile in xcode (#2311 )	2024-08-05 09:48:26 +03:00
Georgi Gerganov	6739eb83c3	whisper : handle empty mel (#2324 )	2024-07-27 20:35:04 +03:00
Matt Stephenson	f68298ce06	whisper : use vulkan as gpu backend when available (#2302 ) * ggml: use vulkan as gpu backend when available Signed-off-by: Matt Stephenson <mstephenson6@users.noreply.github.com> * whisper: enable using vk as default buffer type Signed-off-by: Matt Stephenson <mstephenson6@users.noreply.github.com> --------- Signed-off-by: Matt Stephenson <mstephenson6@users.noreply.github.com>	2024-07-16 10:21:09 +03:00
arizhih	7ae885c1ef	whisper : fix DTW assert (#2299 )	2024-07-15 15:50:36 +03:00
Georgi Gerganov	d207c68822	cmake : use WHISPER_EXTRA_FLAGS (#2294 )	2024-07-09 18:54:18 +03:00
Borislav Stanimirov	16d72504fe	cmake : allow external ggml	2024-07-09 11:38:15 +03:00
Georgi Gerganov	1c31f9d4a8	cmake : try to fix openvino build (#2281 )	2024-07-08 15:36:51 +03:00
Georgi Gerganov	8ecb2f1f68	cmake : remove install of llama convert script [no ci] (#2266 )	2024-07-08 14:53:55 +03:00
Georgi Gerganov	5226c3d45c	make : remove llama prints [no ci] (#2265 )	2024-07-08 14:53:55 +03:00
Georgi Gerganov	dbf9c15e30	talk-llama : sync llama.cpp	2024-07-08 14:53:55 +03:00
Georgi Gerganov	d3f6c34976	examples : fix compile warnings [no ci] (#0 )	2024-07-08 14:53:55 +03:00
Georgi Gerganov	425e2910a3	sync : ggml	2024-07-08 14:53:55 +03:00
Georgi Gerganov	49868aa851	ggml : sync sycl (skip) (#0 )	2024-07-08 14:53:55 +03:00
Georgi Gerganov	ff08e30ab5	scripts : fix sync scripts	2024-07-08 14:53:55 +03:00
Daniel Bevenius	95f2a191c0	ggml : remove unnecessary UNUSED macro call (ggml/880) This commit removes an UNUSED macro call that is not needed as the variable n0 is used in the code and will not produce a warning. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2024-07-08 14:53:55 +03:00
Natsu	00422ec3cf	cmake : add GGML_BUILD and GGML_SHARED macro definitions (llama/8281)	2024-07-08 14:53:55 +03:00
Ouadie EL FAROUKI	c5b05321e9	Enabled more data types for oneMKL gemm_batch (llama/8236)	2024-07-08 14:53:55 +03:00
Johannes Gäßler	5dc636a65a	CUDA: MMQ support for iq4_nl, iq4_xs (llama/8278)	2024-07-08 14:53:55 +03:00
Daniele	73703a144f	CUDA: revert part of the RDNA1 optimizations (llama/8309) The change on the launch_bounds was causing a small performance drop in perplexity of 25 t/s	2024-07-08 14:53:55 +03:00

... 6 7 8 9 10 ...

1844 Commits