Commit Graph

  • defe731263 llamafile: use member variable instead of constant for iq4nlt (llama/11780) Jeffrey Morgan 2025-02-13 09:05:04 -08:00
  • 4e07957bf9 musa: bump MUSA SDK version to rc3.1.1 (llama/11822) R0CKSTAR 2025-02-13 20:28:18 +08:00
  • d2c5154bb5 ggml-cpu : add chunking support to mul_mat_id (llama/11666) Diego Devesa 2025-02-13 01:02:38 +01:00
  • 4fac43fe00 ggml : x2 speed for WASM by optimizing SIMD (llama/11453) Xuan-Son Nguyen 2025-02-13 00:33:45 +01:00
  • 3be9670f17 HIP: Remove GCN from list of devices that avoid MMQ (llama/11831) uvos 2025-02-12 22:25:28 +01:00
  • 86729fcd6d HIP: Switch to std::vector in rocblas version check (llama/11820) uvos 2025-02-12 17:25:03 +01:00
  • 7fbca6304e cleanup: fix compile warnings associated with gnu_printf (llama/11811) bandoti 2025-02-12 10:06:53 -04:00
  • d597f83e1a ggml : fix multi-threaded clamp_f32 (llama/11824) Richard 2025-02-12 13:57:33 +00:00
  • e5edcc6259 ggml-cpu: Fix duplicate MATMUL_INT8 (llama/11817) Weizhao Ouyang 2025-02-12 20:22:58 +08:00
  • 556f773d53 CUDA: fix CUDART_VERSION checks (llama/11821) Johannes Gäßler 2025-02-12 13:16:39 +01:00
  • 91d02de332 Fix #11802: Compile bug - RegQueryValueExA changed to RegQueryValueEx (llama/11803) Sheldon Robinson 2025-02-11 10:55:45 -05:00
  • 1b67d72f87 CUDA: use arch list for compatibility check (llama/11775) Johannes Gäßler 2025-02-11 00:17:22 +01:00
  • 14d7c0368d fix: typos in documentation files (llama/11791) Maxim Evtush 2025-02-10 23:21:31 +01:00
  • db6e19188a vulkan: Make Vulkan optional at runtime (ggml/11493). (llama/11494) Danny Milosavljevic 2025-02-10 07:17:21 +01:00
  • b4b063a5c9 vulkan: add environment variable GGML_VK_PREFER_HOST_MEMORY to avoid VRAM allocation (llama/11592) Wagner Bruna 2025-02-10 03:08:22 -03:00
  • 930b739e7a vulkan: account for lookup tables when checking shared memory size (llama/11502) Jeff Bolz 2025-02-09 01:43:51 -06:00
  • 5981352bb5 ggml: Fix data race in ggml threadpool (llama/11736) Karol Kontny 2025-02-08 15:30:53 +01:00
  • 7561da244e CUDA: fix min. version for movmatrix (llama/11751) Johannes Gäßler 2025-02-08 10:46:07 +01:00
  • be83f342fb vulkan: print shared memory size (llama/11719) Jeff Bolz 2025-02-07 04:26:03 -06:00
  • fd369871f7 SYCL: remove XMX info from print devices (llama/11712) Akarshan Biswas 2025-02-07 14:57:53 +05:30
  • bbd8364f5e ggml : optimize and build warning fix for LoongArch (llama/11709) Jinyang He 2025-02-07 15:38:31 +08:00
  • e4102440ef SYCL: Adjust support condition for norm operators (llama/11674) Akarshan Biswas 2025-02-06 17:12:35 +05:30
  • f8242ec483 ggml : fix LoongArch compile error with 128-bit SIMD (llama/11701) junchao-zhao 2025-02-06 17:20:00 +08:00
  • ef51b4cba4 vulkan: optimize coopmat2 iq2/iq3 callbacks (llama/11521) Jeff Bolz 2025-02-06 00:15:30 -06:00
  • 6f08b24146 vulkan: initial support for IQ4_XS quantization (llama/11501) Rémy O 2025-02-06 07:09:59 +01:00
  • 7c165d7fa8 vulkan: use smaller combined allocations to avoid fragmentation (llama/11551) Jeff Bolz 2025-02-06 00:02:18 -06:00
  • 2f0cf44915 metal : avoid breaking build when metal API predates TARGET_OS_VISION (llama/11690) Charles Duffy 2025-02-05 19:52:31 -06:00
  • b9c972fd0d metal : adjust support conditions for norm operators (llama/11671) Georgi Gerganov 2025-02-05 10:57:42 +02:00
  • 01c9aafbfd CUDA: support for mat. mul. with ne03 != ne13 (llama/11656) Johannes Gäßler 2025-02-05 08:58:31 +01:00
  • bae6bbf487 CUDA: non-contiguous (RMS) norm support (llama/11659) Johannes Gäßler 2025-02-04 22:21:42 +01:00
  • c310272fa0 HIP: force max threads per block to be 1024 (llama/11621) fxzjshm 2025-02-05 02:18:38 +08:00
  • bd0b55dbe0 metal : use residency set for other platforms (llama/11648) Jhen-Jie Hong 2025-02-04 19:07:18 +08:00
  • ba4645db2c rpc: fix known RCE in rpc-server (ggml/1103) Patrick Peng 2025-02-06 09:29:13 -05:00
  • dfc6ca62f3 stream : add beam size parameter(#2836) masahji 2025-02-25 01:39:33 -08:00
  • 47e14c0529 whisper : restore big endian support (#2816) Thomas Fitzsimmons 2025-02-25 09:38:13 +00:00
  • d682e15090 Fixes for Windows (#2790) Judd 2025-02-06 15:37:21 +08:00
  • 46d07b9c85 cmake : fix compile assumptions for power9/etc (#2777) midnight 2025-02-05 04:41:10 -08:00
  • 33ea03f131 authors : update Georgi Gerganov 2025-02-04 13:03:40 +02:00
  • dbcc669e1a sync : ggml Georgi Gerganov 2025-02-04 13:03:09 +02:00
  • 16245b35e4 cmake: Add ability to pass in GGML_BUILD_NUMBER (ggml/1096) Christian Kastner 2025-02-04 00:17:15 +01:00
  • 898c0cb9d1 readme : add maintenance roadmap Georgi Gerganov 2025-02-04 10:50:10 +02:00
  • eb9e5032c4 ci : add stalebot Georgi Gerganov 2025-02-04 09:30:08 +02:00
  • cadfc50eab node : add max_len params in node addon (#2760) billyct 2025-02-04 04:49:06 +08:00
  • 3f91832352 talk-llama : sync llama.cpp Georgi Gerganov 2025-02-03 22:42:26 +02:00
  • cff8868b5f coreml : always convert to "neuralnetwork" (#2770) mgrachten 2025-02-03 21:36:32 +01:00
  • 90e3c5fc40 ci : more git Georgi Gerganov 2025-02-03 21:17:33 +02:00
  • e0f4cef867 ci : install git Georgi Gerganov 2025-02-03 20:12:37 +02:00
  • 234460987e ci : use ubuntu-22.04 instead of ubuntu-latest Georgi Gerganov 2025-02-03 19:50:24 +02:00
  • b8ab126343 cmake : sync cmake scripts Georgi Gerganov 2025-02-03 16:24:38 +02:00
  • edc5d9267c sync : ggml Georgi Gerganov 2025-02-03 16:05:34 +02:00
  • 344b98a44f scripts : fix sync paths Georgi Gerganov 2025-02-03 16:05:27 +02:00
  • dbeb7916b8 CUDA: fix Volta FlashAttention logic (llama/11615) Johannes Gäßler 2025-02-03 13:25:56 +01:00
  • fad2806352 HIP: fix flash_attn_stream_k_fixup warning (llama/11604) Johannes Gäßler 2025-02-02 23:48:29 +01:00
  • 9906792ec3 CUDA/HIP: add support for selectable warp size to mmv (llama/11519) uvos 2025-02-02 22:40:09 +01:00
  • c49ee07ff4 HIP: add GGML_CUDA_CC_IS_* for amd familys as increasing cc archtectures for amd gpus are not supersets of eatch other (llama/11601) uvos 2025-02-02 22:08:05 +01:00
  • f8a831779e CUDA: use mma PTX instructions for FlashAttention (llama/11583) Johannes Gäßler 2025-02-02 19:31:09 +01:00
  • 85451e3612 ci: use sccache on windows instead of ccache (llama/11545) Olivier Chafik 2025-01-31 17:12:40 +00:00
  • 43c744ce8b HIP: require at least HIP 5.5 uvos 2025-01-29 19:36:00 +01:00
  • fc2e44490d HIP: Prepare reduction operators for wave 64 uvos 2025-01-29 19:12:42 +01:00
  • f41fdad200 CUDA/HIP: add warp_size to cuda_device_info uvos 2025-01-29 17:46:23 +01:00
  • 80fa576254 vulkan: implement initial support for IQ2 and IQ3 quantizations (llama/11360) Rémy Oudompheng 2025-01-29 18:29:39 +01:00
  • 75e7d0585e vulkan: Catch pipeline creation failure and print an error message (llama/11436) Jeff Bolz 2025-01-29 09:26:50 -06:00
  • 682a6f5f87 HIP: Supress transformation warning in softmax.cu uvos 2025-01-28 23:06:32 +01:00
  • 115716d109 HIP: Only call rocblas_initialize on rocblas versions with the multiple instantation bug (llama/11080) Nikita Sarychev 2025-01-28 07:42:20 -08:00
  • b2cfef655b cmake : don't fail on GGML_CPU=OFF (llama/11457) someone13574 2025-01-28 09:15:34 -05:00
  • 22e3df0afa SYCL : SOFTMAX F16 mask support and other fixes (llama/11261) Akarshan Biswas 2025-01-28 15:26:58 +05:30
  • 028511d349 AMD: parse the architecture as supplied by gcnArchName (llama/11244) Haus1 2025-01-27 08:58:17 -05:00
  • 70c4038842 metal: Handle null returned from MTLCreateSystemDefaultDevice() (llama/11441) Ihar Hrachyshka 2025-01-27 02:41:59 -05:00
  • 8639c003a9 metal : use residency sets (llama/11427) Georgi Gerganov 2025-01-26 20:06:16 +02:00
  • d5d831da65 cmake: add ggml find package (llama/11369) bandoti 2025-01-26 12:07:48 -04:00
  • 7230a6e1c8 vulkan: compile shaders on-demand (llama/11406) Jeff Bolz 2025-01-25 15:29:57 -06:00
  • a160fa0f3a Hip: disable VMM on hip as it seams that it dosent work in some configurations (llama/11420) uvos 2025-01-25 21:01:12 +01:00
  • 0282ad8fd1 hip : Add hipGraph and VMM support to ROCM (llama/11362) uvos 2025-01-25 00:02:23 +01:00
  • 9e467815d4 CUDA: fix FP16 cuBLAS GEMM (llama/11396) Johannes Gäßler 2025-01-24 21:02:43 +01:00
  • 727891d9bf rocBLAS: Avoid fp32->fp16->fp32 conversion on cdna (llama/11356) uvos 2025-01-24 17:50:49 +01:00
  • c262dc80e2 CPU/CUDA: fix (GQA) mul mat back, add CUDA support (llama/11380) Johannes Gäßler 2025-01-24 12:38:31 +01:00
  • 30767b4c4e cmake : avoid -march=native when reproducible build is wanted (llama/11366) Bernhard M. Wiedemann 2025-01-24 12:21:35 +01:00
  • 16eeb31933 Vulkan-run-test: fix mmq_wg_denoms (llama/11343) amd-dwang 2025-01-23 15:14:28 +08:00
  • ba523d5e22 vulkan: sort shaders for more deterministic binary (llama/11315) Jeff Bolz 2025-01-23 01:07:50 -06:00
  • 3736706139 vulkan: fix diag_mask_inf (llama/11323) Jeff Bolz 2025-01-23 01:01:17 -06:00
  • 58640aa456 rpc : better caching of the base buffer pointer (llama/11331) Radoslav Gerganov 2025-01-21 15:06:41 +02:00
  • 5183a05e56 metal : fix out-of-bounds write (llama/11314) Georgi Gerganov 2025-01-21 08:48:13 +02:00
  • 0dcada42d4 vulkan: fix coopmat2 validation failures (llama/11284) Jeff Bolz 2025-01-20 10:38:32 -06:00
  • d507b4cebe SYCL: Introducing memory host pool (llama/11251) Nicolò Scipione 2025-01-19 14:33:34 +01:00
  • 90171055f3 cmake : add sanitizer flags for llama.cpp (llama/11279) Georgi Gerganov 2025-01-18 16:18:15 +02:00
  • 668306ff2b vulkan: fix coopmat2 flash attention for non-contiguous inputs (llama/11281) Jeff Bolz 2025-01-18 02:26:50 -06:00
  • fdc21fc87b rpc : early register backend devices (llama/11262) Radoslav Gerganov 2025-01-17 10:57:09 +02:00
  • 7183a1eb72 vulkan: support copy from f32 to q4_0/q4_1/q5_0/q5_1/q8_0/iq4_nl (llama/11166) Jeff Bolz 2025-01-16 15:47:10 -06:00
  • 09f3c66648 vulkan: optimize coopmat2 q4_k/q5_k dequant functions. (llama/11206) Jeff Bolz 2025-01-16 15:23:49 -06:00
  • 62e2414620 vulkan: optimize coopmat2 q2_k dequant function (llama/11130) Jeff Bolz 2025-01-16 15:16:39 -06:00
  • de49024e49 CUDA: backwards pass for misc. ops, add tests (llama/11257) Johannes Gäßler 2025-01-16 16:43:38 +01:00
  • db6383094c ggml: aarch64: implement SVE kernels for q4_K_q8_K vector dot (llama/11227) fj-y-saito 2025-01-16 18:11:49 +09:00
  • 164f13c6a9 vulkan: scale caching for k quants + misc fixes (llama/11081) Eve 2025-01-15 19:50:13 +00:00
  • 02aa86230a fix: ggml: fix vulkan-shaders-gen build (llama/10448) Junil Kim 2025-01-15 22:17:42 +09:00
  • 54a2ee648f RoPE: fix back, CUDA support for back + noncont. (llama/11240) Johannes Gäßler 2025-01-15 12:51:37 +01:00
  • 9700cfb0a3 SYCL: Add gated linear attention kernel (llama/11175) Akarshan Biswas 2025-01-15 08:50:17 +05:30
  • 8e0143e205 ggml : add option to not print stack on abort (ggml/1081) William Tambellini 2025-01-23 11:59:08 -08:00
  • f12559d590 ggml-cpu : fix ggml_graph_compute_thread did not terminate on abort. (ggml/1065) issixx 2025-01-17 21:29:08 +09:00
  • 589b40810a ci : dummy commit to trigger CI Georgi Gerganov 2025-02-03 16:32:48 +02:00
  • 7ffcd05267 ruby : Make context accept initial parameters, API to retrieve a segment and more (#2749) KITAITI Makoto 2025-01-21 16:39:54 +09:00