Commit Graph

  • 479bd77169 vulkan: request round-to-even for fp16 in im2col/rope_head (llama/10767) Jeff Bolz 2024-12-10 14:23:17 -06:00
  • d8bf63a41b vulkan: dynamic subgroup size for the remaining k quants (llama/10745) Eve 2024-12-10 19:33:23 +00:00
  • b82c8d76dc CUDA: rename macros to avoid conflicts with WinAPI (llama/10736) Andreas Kieslinger 2024-12-10 18:23:24 +01:00
  • 86346f811e vulkan: disable spirv-opt for coopmat shaders (llama/10763) Jeff Bolz 2024-12-10 11:22:20 -06:00
  • c635f40a34 ggml : remove return from ggml_gallocr_allocate_node (ggml/1048) Daniel Bevenius 2024-12-14 03:23:08 +01:00
  • e0be0de1ee ggml : add check for grad_accs (ggml/1046) Daniel Bevenius 2024-12-13 08:19:38 +01:00
  • 60dc6d003f common : remove old types Georgi Gerganov 2024-12-10 17:19:09 +02:00
  • eb27e0d834 CUDA: fix shared memory access condition for mmv (llama/10740) Johannes Gäßler 2024-12-09 20:07:12 +01:00
  • a682fdce0c vulkan: fix compile warnings (llama/10731) Jeff Bolz 2024-12-09 01:24:01 -06:00
  • 9ffbd3d969 Vulkan: fix NaN in tanh.comp with AMD proprietary driver on Windows (llama/10723) stduhpf 2024-12-08 19:19:19 +01:00
  • 6585a890b4 vulkan: compile a test shader in cmake to check for coopmat2 support (llama/10713) Jeff Bolz 2024-12-08 02:05:55 -06:00
  • d0a050b51f ggml : disable iq4_nl interleave size 8 (llama/10709) Georgi Gerganov 2024-12-07 18:38:15 +02:00
  • e990d1b791 ggml : refactor online repacking (llama/10446) Djip007 2024-12-07 13:37:50 +01:00
  • 4a6d52efe6 Vulkan: VK_KHR_cooperative_matrix support to speed up prompt processing (llama/10597) 0cc4m 2024-12-07 10:24:15 +01:00
  • 8b841d430a metal : Extend how Llama.cpp locates metal resources (llama/10676) Robert Ormandi 2024-12-07 01:55:01 -06:00
  • b74b68212a vulkan: Add VK_NV_cooperative_matrix2 support for mul_mat and flash attention (llama/10206) Jeff Bolz 2024-12-05 13:15:05 -06:00
  • 3a27b2b91b ruby : Add no_speech_thold (#2641) KITAITI Makoto 2024-12-18 18:00:50 +09:00
  • d34445e960 stream : improve consistency in README (#2642) crummyh 2024-12-18 00:43:48 -06:00
  • f897eb7670 whisper : support no_speech_thold (#2625) Karthick 2024-12-17 22:45:47 +05:30
  • 2f2841bfce whisper : add single-timestamp logic (#2629) Karthick 2024-12-17 22:37:08 +05:30
  • 09a1b61218 readme : fix typo (#2637) crummyh 2024-12-17 11:05:35 -06:00
  • 94e7da1ff2 cmake : fix "amd64" processor string (#2638) Georgi Gerganov 2024-12-17 18:34:32 +02:00
  • c4aed6831e vulkan : fix soft_max.comp division by zero (#2633) gn64 2024-12-16 19:34:38 +09:00
  • 199579652e common : add cstdio header Georgi Gerganov 2024-12-16 08:57:04 +02:00
  • d17e7139d8 stream : update build instructions Georgi Gerganov 2024-12-15 21:55:36 +02:00
  • 6a52eaea74 android : fix build and ci (#2624) Thamster 2024-12-14 10:25:53 -05:00
  • 6aa1d7b892 models : fix typo in download-ggml-model.sh (#2623) Michael Rienstra 2024-12-12 08:02:00 -08:00
  • 262e865a70 ruby : Sync whisper.cpp and model download feature (#2617) KITAITI Makoto 2024-12-09 20:17:50 +09:00
  • ed733e85a1 scripts : update to new build system Georgi Gerganov 2024-12-09 11:30:16 +02:00
  • 5980b1ae77 devops : add cmake Georgi Gerganov 2024-12-08 23:09:26 +02:00
  • 0415a66044 devops : update make commands Georgi Gerganov 2024-12-08 23:07:29 +02:00
  • 7d134e3737 ggml : remove old files (skip) (#0) Georgi Gerganov 2024-12-08 23:04:26 +02:00
  • 9df53b357e ggml : sync remnants (skip) (#0) Georgi Gerganov 2024-12-08 22:48:25 +02:00
  • b2115b4d9b scripts : remove amx from sync Georgi Gerganov 2024-12-08 22:48:14 +02:00
  • 0164427dd5 ci : disable freeBSD builds [no ci] Georgi Gerganov 2024-12-08 15:52:57 +02:00
  • 627b11c78a readme : update build instructions Georgi Gerganov 2024-12-08 15:48:14 +02:00
  • 472464453d ci : disable CUDA and Android builds Georgi Gerganov 2024-12-08 15:36:01 +02:00
  • 11dddfbc9e ci : disable Obj-C build + fixes Georgi Gerganov 2024-12-08 13:35:35 +02:00
  • 384e214cc7 make : shim cmake Georgi Gerganov 2024-12-06 15:34:53 +02:00
  • f2c680f893 talk-llama : sync llama.cpp Georgi Gerganov 2024-12-05 14:30:33 +02:00
  • fbe66da0e5 sync : ggml Georgi Gerganov 2024-12-05 14:29:18 +02:00
  • a815940e0e ggml : add predefined list of CPU backend variants to build (llama/10626) Diego Devesa 2024-12-04 14:45:40 +01:00
  • 904e307bce ggml-cpu : fix HWCAP2_I8MM value (llama/10646) Diego Devesa 2024-12-04 14:40:44 +01:00
  • 491ec076b4 vulkan: Implement "fast divide" (mul+shift) for unary ops like copy (llama/10642) Jeff Bolz 2024-12-04 01:28:59 -06:00
  • 966433fdf2 SYCL : Move to compile time oneMKL interface backend selection for NVIDIA backend (llama/10584) Nicolò Scipione 2024-12-04 02:29:20 +01:00
  • 6f1ba9d82d Avoid using __fp16 on ARM with old nvcc (llama/10616) Frankie Robertson 2024-12-04 02:41:37 +02:00
  • 015ecd0001 vulkan: optimize and reenable split_k (llama/10637) Jeff Bolz 2024-12-03 13:29:54 -06:00
  • b7c64a4352 ggml: add GGML_SET Metal kernel + i32 CPU kernel (ggml/1037) PAB 2024-12-04 09:19:30 +01:00
  • 7895d39508 ggml : add GGML_PAD_REFLECT_1D operation (ggml/1034) PAB 2024-12-03 20:20:04 +01:00
  • 22616f00f9 files : remove make artifacts Georgi Gerganov 2024-12-03 20:29:32 +02:00
  • 02c6fcbc2c common : fix compile warning Georgi Gerganov 2024-12-03 20:25:37 +02:00
  • 3daeacad24 ggml : move AMX to the CPU backend (llama/10570) Diego Devesa 2024-12-03 20:22:12 +02:00
  • 4d73962da4 metal : small-batch mat-mul kernels (llama/10581) Georgi Gerganov 2024-12-03 11:52:33 +02:00
  • 068812650e SYCL: Fix and switch to GGML_LOG system instead of fprintf (llama/10579) Akarshan Biswas 2024-12-02 12:34:11 +05:30
  • 4b7e059e15 ggml-cpu: replace AArch64 NEON assembly with intrinsics in ggml_gemv_q4_0_4x4_q8_0() (llama/10567) Adrien Gallouët 2024-11-30 18:13:18 +01:00
  • 30e35d7271 vulkan: Dynamic subgroup size support for Q6_K mat_vec (llama/10536) Eve 2024-11-30 07:00:02 +00:00
  • 3623bd58f2 ggml : fix I8MM Q4_1 scaling factor conversion (llama/10562) Georgi Gerganov 2024-11-29 16:25:39 +02:00
  • cb847c20a7 ggml-cpu: fix typo in gemv/gemm iq4_nl_4_4 (llama/10580) Shupei Fan 2024-11-29 21:49:02 +08:00
  • 964b154a2a sycl : offload of get_rows set to 0 (llama/10432) Alberto Cabrera Pérez 2024-11-29 12:38:45 +00:00
  • d7c2a04bce sycl : Reroute permuted mul_mats through oneMKL (llama/10408) Alberto Cabrera Pérez 2024-11-29 09:49:43 +00:00
  • 2bb4ca9cba CANN: RoPE operator optimization (llama/10563) Chenguang Li 2024-11-29 14:46:55 +08:00
  • a753a82462 vulkan: get the first command buffer submitted sooner (llama/10499) Jeff Bolz 2024-11-29 00:18:02 -06:00
  • 276b08d8f0 ggml : remove redundant copyright notice + update authors Georgi Gerganov 2024-11-28 20:46:40 +02:00
  • 4ca1e72fe0 ggml : fix row condition for i8mm kernels (llama/10561) Georgi Gerganov 2024-11-28 14:56:37 +02:00
  • 16a66f103f cmake : fix ARM feature detection (llama/10543) Georgi Gerganov 2024-11-28 14:56:23 +02:00
  • 330273901f ggml-cpu: support IQ4_NL_4_4 by runtime repack (llama/10541) Shupei Fan 2024-11-28 20:52:03 +08:00
  • 42099a9342 kompute : improve backend to pass test_backend_ops (llama/10542) Sergio López 2024-11-28 12:51:38 +01:00
  • 90dd5fca9c CANN: Fix SOC_TYPE compile bug (llama/10519) leo-pony 2024-11-28 15:25:24 +08:00
  • 2490f2a7f8 CANN: ROPE operator optimization (llama/10540) Chenguang Li 2024-11-28 14:24:46 +08:00
  • 230e985633 Add some minimal optimizations for CDNA (llama/10498) uvos 2024-11-27 17:10:08 +01:00
  • ae24083f23 metal : fix group_norm support condition (llama/0) Georgi Gerganov 2024-11-27 11:22:14 +02:00
  • 6463e36369 vulkan: define all quant data structures in types.comp (llama/10440) Jeff Bolz 2024-11-27 01:32:54 -06:00
  • b3301f7d82 vulkan: Handle GPUs with less shared memory (llama/10468) Jeff Bolz 2024-11-27 01:30:27 -06:00
  • ab5d4d93ec vulkan: further optimize q5_k mul_mat_vec (llama/10479) Jeff Bolz 2024-11-27 01:21:59 -06:00
  • 2d6e9dd723 vulkan: skip integer div/mod in get_offsets for batch_idx==0 (llama/10506) Jeff Bolz 2024-11-27 01:08:54 -06:00
  • 2f16e51553 vulkan: optimize Q2_K and Q3_K mul_mat_vec (llama/10459) Jeff Bolz 2024-11-27 01:00:50 -06:00
  • 0f0994902f mtgpu: Add MUSA_DOCKER_ARCH in Dockerfiles && update cmake and make (llama/10516) R0CKSTAR 2024-11-27 00:00:41 +08:00
  • 5e1fcc1780 vulkan: fix group_norm (llama/10496) Jeff Bolz 2024-11-26 09:45:05 -06:00
  • 48f421de23 cmake : enable warnings in llama (llama/10474) Georgi Gerganov 2024-11-26 14:18:08 +02:00
  • e7afb2b991 ggml-cpu: cmake add arm64 cpu feature check for macos (llama/10487) Charles Xu 2024-11-26 12:37:05 +01:00
  • 9a5ef7b169 CANN: Improve the Inferencing Performance for Ascend NPU Device (llama/10454) Shanshan Shen 2024-11-26 18:08:37 +08:00
  • 453cc0fcf1 CANN: RoPE and CANCAT operator optimization (llama/10488) Chenguang Li 2024-11-26 17:31:05 +08:00
  • 78dfec6bc5 vulkan: Fix a vulkan-shaders-gen arugment parsing error (llama/10484) Junil Kim 2024-11-26 10:47:20 +09:00
  • f6d518fc4c metal : enable mat-vec kernels for bs <= 4 (llama/10491) Georgi Gerganov 2024-11-25 21:49:31 +02:00
  • ac33379a35 llama : accept a list of devices to use to offload a model (llama/10497) Diego Devesa 2024-11-25 19:30:06 +01:00
  • 77e3e4a090 ggml : add support for dynamic loading of backends (llama/10469) Diego Devesa 2024-11-25 15:13:39 +01:00
  • b840bb09be metal : minor code formatting Georgi Gerganov 2024-11-25 15:08:04 +02:00
  • 8b1c1c30a7 ggml : do not use ARM features not included in the build (llama/10457) Diego Devesa 2024-11-23 14:41:12 +01:00
  • 4b81335f75 CANN: Support Ascend310P to accelerate F32 and F16 Model (llama/10216) leo-pony 2024-11-22 14:07:20 +08:00
  • 2a4b5c9d7e cuda : optimize argmax (llama/10441) Diego Devesa 2024-11-21 18:18:50 +01:00
  • 04662748aa vulkan: predicate max operation in soft_max shaders/soft_max (llama/10437) Jeff Bolz 2024-11-20 13:47:36 -06:00
  • a117279e13 vulkan: copy iq4_nl LUT into shared memory (llama/10409) Jeff Bolz 2024-11-20 01:40:18 -06:00
  • bbb292ed38 vulkan: further optimize mul_mat_vec using larger loads (llama/10387) Jeff Bolz 2024-11-20 01:11:00 -06:00
  • 95e8901e71 add cmake rvv support (llama/10411) haopeng 2024-11-20 04:10:31 +08:00
  • 4af9626702 CUDA: remove unnecessary warp reduce in FA (ggml/1032) mahorozte 2024-12-03 21:11:43 +08:00
  • c52d1035de feat: add GGML_UNARY_OP_ARGMAX Metal kernel (ggml/1019) PAB 2024-12-02 19:27:24 +01:00
  • 5773a14980 metal : add GGML_OP_CONV_TRANSPOSE_1D kernels (ggml/1026) PAB 2024-11-28 09:25:06 +01:00
  • 6939147c47 Do not include arm_neon.h when compiling CUDA code (ggml/1028) Frankie Robertson 2024-11-26 15:50:26 +02:00
  • 98f9916c9f ggml-opt: fix data corruption (ggml/1022) Johannes Gäßler 2024-11-20 14:56:04 +01:00
  • 021eef1000 ruby : Add low-level methods to transcribe (#2585) KITAITI Makoto 2024-11-28 17:33:07 +09:00