Commit Graph

  • d1286cf32b ggml : support bcast ggml_soft_max_ext, ggml_flash_attn_ext (llama/14435) Georgi Gerganov 2025-07-12 14:33:49 +03:00
  • 2e04b81f3e opencl : fix possible buffer overflow in dump_tensor (llama/14490) zhouwg 2025-07-02 20:38:10 +08:00
  • cd87a2f7e0 opencl : skip empty nodes on cgraph compute (llama/14491) Eric Zhang 2025-07-02 19:00:04 +08:00
  • e43c38f9f1 opencl : update upscale to support align corners (llama/14488) lhez 2025-07-02 00:07:42 -07:00
  • ab850d4680 ggml : Callback before abort (llama/14481) Björn Ganster 2025-07-02 07:19:31 +02:00
  • cdf5e72163 ci : disable fast-math for Metal GHA CI (llama/14478) Georgi Gerganov 2025-07-01 18:04:08 +03:00
  • 32d7c10766 CANN: update aclnnGroupedMatmulV2 to aclnnGroupedMatmulV3 (llama/14411) Chenguang Li 2025-07-01 16:47:30 +08:00
  • 3c7939cfe5 vulkan: Split large mul_mat_id to fit in shared memory (llama/14451) Jeff Bolz 2025-07-01 03:43:08 -05:00
  • 6fc80e8456 add GELU_ERF (llama/14455) Sigbjørn Skjæret 2025-07-01 10:14:21 +02:00
  • 19b9aaf044 vulkan : implement bilinear interpolation for ggml_upscale/ggml_interpolate (ggml/1291) Acly 2025-07-03 19:58:12 +02:00
  • f98cb6607b vulkan : implement ggml_roll (ggml/1290) Acly 2025-07-03 19:47:15 +02:00
  • 5ea5c58768 ggml : add version function to get lib version (ggml/1286) Daniel Bevenius 2025-07-02 13:55:32 +02:00
  • 869335f2d5 server : add dtw.params for v3-large-turbo (#3307) accessiblepixel 2025-07-07 10:51:15 +01:00
  • d9999d54c8 feat: support vad for addon.node (#3301) Lin Xiaodong 2025-07-02 18:14:29 +08:00
  • bca021c974 sync : ggml Georgi Gerganov 2025-07-01 12:21:19 +03:00
  • 1f816de7da talk-llama : sync llama.cpp Georgi Gerganov 2025-07-01 12:21:09 +03:00
  • c4ea72be9a ggml : remove trailing whitespace (llama/0) Georgi Gerganov 2025-07-01 11:05:48 +03:00
  • 1e930ab1b8 opencl : add GEGLU, REGLU, SWIGLU (llama/14456) lhez 2025-07-01 00:19:16 -07:00
  • b5b237d49a Add Conv2d for CPU (llama/14388) Aman Gupta 2025-06-30 23:57:04 +08:00
  • 679f31a9d1 metal : disable fast-math for some cpy kernels (llama/14460) Georgi Gerganov 2025-06-30 17:04:05 +03:00
  • e29e36aee7 ggml-cpu: sycl: Re-enable exp f16 (llama/14462) Romain Biessy 2025-06-30 14:52:02 +02:00
  • 6bb1234a56 cmake : Remove redundant include path in CMakeLists.txt (llama/14452) xiaobing318 2025-06-30 17:48:24 +08:00
  • 3239359bd1 scripts : make the shell scripts cross-platform (llama/14341) Vedran Miletić 2025-06-30 10:17:18 +02:00
  • e81be92931 SYCL: disable faulty fp16 exp kernel (llama/14395) Akarshan Biswas 2025-06-29 21:07:58 +05:30
  • 130044f228 ggml : fix unmerged GGML_FPxx_TO_FPxx refactoring (llama/14443) Sigbjørn Skjæret 2025-06-29 14:38:10 +02:00
  • 8bc638ee56 ggml : implement REGLU/GEGLU/SWIGLU ops (llama/14158) Sigbjørn Skjæret 2025-06-29 11:04:10 +02:00
  • 00b36237ba vulkan: Add fusion support for RMS_NORM+MUL (llama/14366) Jeff Bolz 2025-06-29 02:43:36 -05:00
  • b900ee424c CUDA: add bf16 and f32 support to cublas_mul_mat_batched (llama/14361) Aman Gupta 2025-06-29 01:30:53 +08:00
  • f641a4c410 vulkan: handle noncontig in the final case of ggml_vk_get_cpy_pipeline (llama/14378) Jeff Bolz 2025-06-28 10:36:40 -05:00
  • 9e48afba2f vulkan: lock accesses of pinned_memory vector (llama/14333) Jeff Bolz 2025-06-28 10:17:09 -05:00
  • f31ed384f4 fix async_mode bug (llama/14432) Xinpeng Dou 2025-06-28 17:35:41 +08:00
  • 0b09f5bbad vulkan: Fix GGML_VULKAN_SHADER_DEBUG_INFO (llama/14427) Jeff Bolz 2025-06-27 22:35:30 -05:00
  • 48fb51f314 ggml : add ggml_set_rows (llama/14274) Radoslav Gerganov 2025-06-27 16:41:40 +03:00
  • 566462a5c0 cmake: regen vulkan shaders when shaders-gen sources change (llama/14398) bandoti 2025-06-26 13:46:53 -03:00
  • c300f1e32d metal : add special-case mat-vec mul for ne00 == 4 (llama/14385) Georgi Gerganov 2025-06-26 15:51:19 +03:00
  • c848b9fbef metal : batch rows copy in a single threadgroup (llama/14384) Georgi Gerganov 2025-06-26 15:50:15 +03:00
  • a5e6a3c953 musa: enable fp16 mma (all) and cublas on qy2 (llama/13842) R0CKSTAR 2025-06-26 12:11:59 +08:00
  • 16aa7d151d ggml-cpu: enable IBM NNPA Vector Intrinsics (llama/14317) Aaron Teo 2025-06-26 05:49:04 +08:00
  • 99764f5767 ggml : do not output unprintable characters on GGUF load failure (llama/14381) Sigbjørn Skjæret 2025-06-25 23:26:51 +02:00
  • fc28594112 sycl: GGML_SYCL_DISABLE_OPT on by default for all Intel Devices (llama/13973) Anton Mitkov 2025-06-25 17:09:55 +01:00
  • acfbf2921b opencl: ref count ggml_backend_opencl_context and refactor profiling (llama/14254) lhez 2025-06-24 11:46:25 -07:00
  • 6a1d12a8ea CUDA/HIP: optimize mmv paths taken for HIP devices (llama/14324) uvos 2025-06-24 01:12:56 +02:00
  • 06b01ba87b CUDA: mul_mat_v support for batch sizes > 1 (llama/14262) Johannes Gäßler 2025-06-23 13:11:31 +02:00
  • 791201a974 HIP: enable vec fattn on RDNA4 (llama/14323) uvos 2025-06-22 16:51:23 +02:00
  • abb650c0ec CUDA: add mean operation (llama/14313) Aman Gupta 2025-06-22 12:39:54 +08:00
  • e036676795 Add support for VK_EXT_debug_utils to add labels to Vulkan objects. (llama/13792) Markus Tavenrath 2025-06-21 08:17:12 +02:00
  • c1418b9906 metal : fix thread-safety (llama/14300) Georgi Gerganov 2025-06-21 08:04:18 +03:00
  • 9d7cb80f04 ggml-cpu : "align corners" for bilinear upscale/downscale (ggml/1285) Acly 2025-07-01 09:11:00 +02:00
  • 515df20351 ggml-quants : rename best_mad to best_error (ggml/1283) Daniel Bevenius 2025-06-24 06:10:16 +02:00
  • c88ffbf9ba ci : use selective copy for musa image (#3296) Daniel Bevenius 2025-06-27 15:43:56 +02:00
  • 7069394447 ci: set fail-fast to false in docker.yml (#3294) Daniel Bevenius 2025-06-27 09:55:56 +02:00
  • f8abbeb234 ruby : add Whisper::VERSION (#3292) KITAITI Makoto 2025-06-27 11:41:26 +09:00
  • 32cf4e2aba whisper : add version function (#3289) Daniel Bevenius 2025-06-26 18:09:42 +02:00
  • 35034c5aea ci : add should_release variable (#3288) Daniel Bevenius 2025-06-26 16:29:29 +02:00
  • 897b071dc6 docs : add cmake "-j" flag in README.md (#3284) toboil-features 2025-06-26 14:23:19 +03:00
  • 4daf7050ca ci : add support for tag-based releases (#3287) Daniel Bevenius 2025-06-25 21:43:58 +02:00
  • a8d002cfd8 release : v1.7.6 Georgi Gerganov 2025-06-25 16:47:03 +03:00
  • 06bdaa6c0c bench : update benches Georgi Gerganov 2025-06-25 16:45:19 +03:00
  • dc8dda60ee bench : print system info before ctx check Georgi Gerganov 2025-06-25 15:59:23 +03:00
  • 1ad258ca31 stream : add nullptr check of whisper_context (#3283) Daniel Bevenius 2025-06-25 14:16:31 +02:00
  • 7dd2997a01 ci : enable main-cuda build (#3282) Daniel Bevenius 2025-06-25 12:12:36 +02:00
  • c85b1ae84e bindings.java : update java example (#3281) Joas Dev 2025-06-24 23:35:38 -05:00
  • 0083335ba0 coreml : backport CoreML features to macos < 14 (#3255) glaszig 2025-06-24 04:24:27 -03:00
  • 9c47902308 ci : reduce musa image size (#3277) Daniel Bevenius 2025-06-24 08:20:28 +02:00
  • a0d2c632e4 whisper : add .gitignore entries for OpenVINO support (#3276) Yukimasa Funaoka 2025-06-24 14:50:16 +09:00
  • 4d6ae52ed3 command: output commands to text file (#3273) Aaron Ang 2025-06-23 21:41:21 -07:00
  • a422176937 ci : add apt-get clean to musa Dockerfile (#3275) Daniel Bevenius 2025-06-23 12:34:44 +02:00
  • cead8f5357 ruby : specify Apple frameworks explicitly on build (#3270) KITAITI Makoto 2025-06-23 13:34:05 +09:00
  • e6c10cf3d5 talk-llama : sync llama.cpp Georgi Gerganov 2025-06-20 21:18:44 +03:00
  • d65a579a0a sync : ggml Georgi Gerganov 2025-06-20 21:16:19 +03:00
  • b68222f92c CUDA: add conv_2d_transpose (llama/14287) Aman Gupta 2025-06-20 22:48:24 +08:00
  • a455dcb04c sycl: add usage of enqueue_functions extension (llama/14244) Nicolò Scipione 2025-06-20 15:07:21 +02:00
  • af7168174c Implement GGML_CPU_ALL_VARIANTS for PowerPC (llama/14286) Christian Kastner 2025-06-20 12:17:32 +00:00
  • 33d1f0a3e0 cuda : synchronize graph capture and cublas handle destruction (llama/14288) Diego Devesa 2025-06-20 04:57:36 -07:00
  • 018b2d340e ggml : fix repack work size for mul_mat_id (llama/14292) Georgi Gerganov 2025-06-20 11:19:15 +03:00
  • 694f435d22 ggml: Update KleidiAI to v1.9.0 (llama/14277) Charles Xu 2025-06-20 09:51:01 +02:00
  • 5efd43c956 CUDA: add conv_2d_dw (llama/14265) Aman Gupta 2025-06-20 09:50:24 +08:00
  • 71adde9203 ggml-cpu : remove unnecesary arm feature detection (llama/14281) Diego Devesa 2025-06-19 12:24:14 -07:00
  • cef59c1e26 build : suppress gcc15 compile warnings (llama/14261) fanyang 2025-06-19 20:49:48 +08:00
  • a02a2d4240 sycl: Cleanup codepaths in Get Rows in sycl backend (llama/14215) Anton Mitkov 2025-06-19 11:40:21 +01:00
  • be4ea0826b llamafile : support s390x SIMD instruction set (llama/14273) Aaron Teo 2025-06-19 17:48:54 +08:00
  • 1aca7b5c8a Vulkan: Set device max size for host memory to avoid OOM warning and fallback to CPU buffer (llama/14249) 0cc4m 2025-06-19 09:15:42 +02:00
  • b251d739ad metal : add mean kernel (llama/14267) Georgi Gerganov 2025-06-19 08:05:21 +03:00
  • 203451bcba ggml-cpu: reduce asm calls for hsum (llama/14037) Aaron Teo 2025-06-19 01:10:08 +08:00
  • 34940abe53 ggml-cpu: fix uncaught underscore terminators (llama/14023) Aaron Teo 2025-06-19 01:06:49 +08:00
  • 4fc9c34126 ggml: Add Apple support for GGML_CPU_ALL_VARIANTS (llama/14258) Charles Xu 2025-06-18 13:40:07 +02:00
  • 471df139fa Add ggml_roll (ggml/1274) Acly 2025-06-18 13:34:50 +02:00
  • 3e65f518dd android : update CMakeLists.txt to use FetchContent for ggml (#3268) Daniel Bevenius 2025-06-19 16:06:42 +02:00
  • 17bece1885 cmake : fix android build (#3265) Georgi Gerganov 2025-06-19 09:24:41 +03:00
  • ecb8f3c2b4 examples : add stereo to mono conversion in read_audio_data (#3266) Daniel Bevenius 2025-06-18 17:41:43 +02:00
  • 2f60ebc3c2 talk-llama : sync llama.cpp Georgi Gerganov 2025-06-18 10:22:47 +03:00
  • 69061e356f sync : ggml Georgi Gerganov 2025-06-18 10:22:11 +03:00
  • 0e068779c7 cmake: remove shader-gen step-targets from ggml-vulkan (llama/14226) bandoti 2025-06-17 17:33:25 -03:00
  • ac8a303c9a ggml-cpu : remove the weak alias trick (llama/14221) xctan 2025-06-17 17:58:32 +08:00
  • 2a84593960 musa: fix build warning (unused variable) (llama/14231) R0CKSTAR 2025-06-17 17:48:08 +08:00
  • 44871c8a3e llama : add thread safety test (llama/14035) Diego Devesa 2025-06-16 08:11:43 -07:00
  • ad6cd94a3a cmake: clean up external project logic for vulkan-shaders-gen (llama/14179) bandoti 2025-06-16 10:32:13 -03:00
  • dbad9d8fba HIP: disable rocwmma on gfx12 by default until rocm 7.0 (llama/14202) uvos 2025-06-16 13:47:38 +02:00
  • 518835ee56 ggml: Add Android support for GGML_CPU_ALL_VARIANTS (llama/14206) Charles Xu 2025-06-16 11:47:57 +02:00
  • a3d1c55c66 vulkan: mutex around vkQueueSubmit (llama/14127) Jeff Bolz 2025-06-16 00:21:08 -06:00