llama-cpp-turboquant

History

Gabe Goodhart bde188d60f metal: TRI, FILL, EXPM1, SOFTPLUS (#16623 ) * feat(wip): Port initial TRI impl from pervious work The kernel does not work and is not optimized, but the code compiles and runs, so this will be the starting point now that the core op has been merged. Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix: Remove argument for constant val override This was added in the original draft, but later removed. With this, the kernel now passes tests. Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Move the ttype conditional to templating to avoid conditional in kernel Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix: Type fixes Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * feat: Add softplus for metal Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Add EXPM1 for metal Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * feat: Add FILL for metal Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * refactor: Branchless version of tri using _ggml_vec_tri_cmp as a mask Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * fix: Remove unused arguments Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> * refactor: Use select instead of branch for softplus non-vec Branch: ggml-cumsum-tri Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> --------- Signed-off-by: Gabe Goodhart <ghart@us.ibm.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>		2025-12-04 19:12:19 +02:00
..
ggml-blas	sync : whisper.cpp (ggml/1359)	2025-09-29 17:43:58 +03:00
ggml-cann	CANN: Disable Ger operator of OUT_PROD on 310p device (#17563 )	2025-12-02 20:35:23 +08:00
ggml-cpu	ggml-cpu : remove asserts always evaluating to false (#17728 )	2025-12-04 13:16:38 +01:00
ggml-cuda	CUDA: generalized (mma) FA, add Volta support (#17505 )	2025-12-03 16:57:05 +01:00
ggml-hexagon	hexagon: add support for ROPE_NEOX (#17458 )	2025-11-23 18:55:56 -08:00
ggml-hip	HIP: fix AMDGPU_TARGETS, update documentation (#16803 )	2025-10-27 21:39:49 +01:00
ggml-metal	metal: TRI, FILL, EXPM1, SOFTPLUS (#16623 )	2025-12-04 19:12:19 +02:00
ggml-musa	CUDA: faster tile FA, add oob checks, more HSs (#16492 )	2025-10-11 20:54:32 +02:00
ggml-opencl	model: LFM2-VL fixes (#17577 )	2025-11-30 21:57:31 +01:00
ggml-rpc	rpc : cache and reuse compute graphs (#15405 )	2025-11-28 08:33:51 +00:00
ggml-sycl	enhance argsort for UT (#17573 )	2025-12-02 08:56:46 +08:00
ggml-vulkan	vulkan: Reduce temporary memory usage for TOP_K (#17623 )	2025-12-02 19:22:04 +01:00
ggml-webgpu	ggml webgpu: add support for emscripten builds (#17184 )	2025-12-03 10:25:34 +01:00
ggml-zdnn	zdnn: refactor codebase + add docs (#16178 )	2025-09-23 14:53:05 +08:00
CMakeLists.txt	build : move _WIN32_WINNT definition to headers (#17736 )	2025-12-04 07:04:02 +01:00
ggml-alloc.c	ggml : add GGML_SCHED_NO_REALLOC option to disable reallocations in ggml_backend_sched (#17276 )	2025-11-28 17:33:23 +02:00
ggml-backend-impl.h	rpc : add support for multiple devices (#16276 )	2025-10-04 12:49:16 +03:00
ggml-backend-reg.cpp	Add experimental ggml-hexagon backend for the Hexagon NPU (#16547 )	2025-10-22 13:47:09 -07:00
ggml-backend.cpp	ggml : remove redundant n_copies check when setting input/output (#17612 )	2025-12-02 12:52:45 +01:00
ggml-common.h	llama : add gpt-oss (#15091 )	2025-08-05 22:10:36 +03:00
ggml-impl.h	ggml : add ops SOFTPLUS, EXPM1, TRI, SOLVE_TRI, CUMSUM (#17063 )	2025-11-13 20:54:47 +02:00
ggml-opt.cpp	finetune: SGD optimizer, more CLI args (#13873 )	2025-08-14 12:03:57 +02:00
ggml-quants.c	ggml : fix uninitialized is_on_grid in quantize_row_iq3_xxs_impl (#15928 )	2025-09-23 10:25:20 +02:00
ggml-quants.h	llama : add gpt-oss (#15091 )	2025-08-05 22:10:36 +03:00
ggml-threading.cpp	ggml : build backends as libraries (#10256 )	2024-11-14 18:04:35 +01:00
ggml-threading.h	remove CMAKE_WINDOWS_EXPORT_ALL_SYMBOLS (#10797 )	2024-12-12 19:02:49 +01:00
ggml.c	model: LFM2-VL fixes (#17577 )	2025-11-30 21:57:31 +01:00
ggml.cpp	ggml : Print backtrace on uncaught C++ exceptions (ggml/1232)	2025-06-01 13:43:57 +03:00
gguf.cpp	ggml, llama : use defaulted constructors/destructors (#17649 )	2025-12-03 07:12:18 +01:00