add amx kernel for gemm (llama/8998)

mirror of https://github.com/ggerganov/whisper.cpp.git synced 2025-06-16 05:48:09 +00:00

add intel amx isa detection

add vnni kernel for gemv cases

add vnni and amx kernel support for block_q8_0

code cleanup

fix packing B issue

enable openmp

fine tune amx kernel

switch to aten parallel pattern

add error message for nested parallelism

code cleanup

add f16 support in ggml-amx

add amx kernels for QK_K quant formats: Q4_K, Q5_K, Q6_K and IQ4_XS

update CMakeList

update README

fix some compilation warning

fix compiler warning when amx is not enabled

minor change

ggml-ci

move ggml_amx_init from ggml.c to ggml-amx/mmq.cpp

ggml-ci

update CMakeLists with -mamx-tile, -mamx-int8 and -mamx-bf16

ggml-ci

add amx as an ggml-backend

update header file, the old path for immintrin.h has changed to ggml-cpu-impl.h

minor change

update CMakeLists.txt

minor change

apply weight prepacking in set_tensor method in ggml-backend

fix compile error

ggml-ci

minor change

ggml-ci

update CMakeLists.txt

ggml-ci

add march dependency

minor change

ggml-ci

change ggml_backend_buffer_is_host to return false for amx backend

ggml-ci

fix supports_op

use device reg for AMX backend

ggml-ci

minor change

ggml-ci

minor change

fix rebase

set .buffer_from_host_ptr to be false for AMX backend

This commit is contained in:

Ma Mingfei

2024-10-18 13:34:36 +08:00

committed by

Georgi Gerganov

parent 28b044dad9

commit e1936eb2a5

5 changed files with 66 additions and 1 deletions

									
										8

ggml/src/ggml.c
									
												View File
												
				@ -23219,6 +23219,14 @@ int ggml_cpu_has_avx512_bf16(void) {

				#endif

				}

				int ggml_cpu_has_amx_int8(void) {

				#if defined(__AMX_INT8__)

				    return 1;

				#else

				    return 0;

				#endif

				}

				int ggml_cpu_has_fma(void) {

				#if defined(__FMA__)

				    return 1;

add amx kernel for gemm (llama/8998)

8 ggml/src/ggml.c Unescape Escape View File

8

ggml/src/ggml.c

View File