Commit Graph

  • 8da6fd4dff vulkan : initialize vk_buffer_struct members to VK_NULL_HANDLE (ggml/893) Tony Wasserka 2024-07-20 20:49:44 +02:00
  • ab8ec9e940 cmake : only enable GGML_NATIVE and x86 flags if not crosscompiling (ggml/885) Borislav Stanimirov 2024-07-12 17:24:20 +03:00
  • 701265bf38 scripts : sync new files (#0) Georgi Gerganov 2024-08-08 14:00:51 +03:00
  • fe36c90971 cmake : fix compile in xcode (#2311) Daven Sanassy 2024-08-05 07:48:26 +01:00
  • 6739eb83c3 whisper : handle empty mel (#2324) Georgi Gerganov 2024-07-27 20:35:04 +03:00
  • f68298ce06 whisper : use vulkan as gpu backend when available (#2302) Matt Stephenson 2024-07-16 03:21:09 -04:00
  • 7ae885c1ef whisper : fix DTW assert (#2299) arizhih 2024-07-15 14:50:36 +02:00
  • d207c68822 cmake : use WHISPER_EXTRA_FLAGS (#2294) Georgi Gerganov 2024-07-09 18:54:18 +03:00
  • 16d72504fe cmake : allow external ggml Borislav Stanimirov 2024-07-08 17:08:55 +03:00
  • 1c31f9d4a8 cmake : try to fix openvino build (#2281) Georgi Gerganov 2024-07-08 15:36:51 +03:00
  • 8ecb2f1f68 cmake : remove install of llama convert script [no ci] (#2266) Georgi Gerganov 2024-07-08 14:21:04 +03:00
  • 5226c3d45c make : remove llama prints [no ci] (#2265) Georgi Gerganov 2024-07-08 14:19:36 +03:00
  • dbf9c15e30 talk-llama : sync llama.cpp Georgi Gerganov 2024-07-08 14:14:17 +03:00
  • d3f6c34976 examples : fix compile warnings [no ci] (#0) Georgi Gerganov 2024-07-08 14:09:09 +03:00
  • 425e2910a3 sync : ggml Georgi Gerganov 2024-07-08 13:50:28 +03:00
  • 49868aa851 ggml : sync sycl (skip) (#0) Georgi Gerganov 2024-07-08 13:50:14 +03:00
  • ff08e30ab5 scripts : fix sync scripts Georgi Gerganov 2024-07-08 13:48:14 +03:00
  • 95f2a191c0 ggml : remove unnecessary UNUSED macro call (ggml/880) Daniel Bevenius 2024-07-08 12:03:42 +02:00
  • 00422ec3cf cmake : add GGML_BUILD and GGML_SHARED macro definitions (llama/8281) Natsu 2024-07-05 22:29:35 +08:00
  • c5b05321e9 Enabled more data types for oneMKL gemm_batch (llama/8236) Ouadie EL FAROUKI 2024-07-05 13:23:25 +01:00
  • 5dc636a65a CUDA: MMQ support for iq4_nl, iq4_xs (llama/8278) Johannes Gäßler 2024-07-05 09:06:31 +02:00
  • 73703a144f CUDA: revert part of the RDNA1 optimizations (llama/8309) Daniele 2024-07-05 07:06:09 +00:00
  • e89fdceec2 CUDA: fix MMQ stream-k rounding if ne00 % 128 != 0 (llama/8311) Johannes Gäßler 2024-07-05 09:05:34 +02:00
  • 29a2739d27 Fix WARP_SIZE=16 bug of Intel GPU (llama/8266) luoyu-intel 2024-07-05 05:06:13 +00:00
  • ee6d17f6b4 rm get_work_group_size() by local cache for performance (llama/8286) Neo Zhang Jianyu 2024-07-05 10:32:29 +08:00
  • 95e90823d9 Define and optimize RDNA1 (llama/8085) Daniele 2024-07-03 23:02:58 +00:00
  • 005cc45df3 fix typo (llama/8267) Judd 2024-07-03 20:40:16 +08:00
  • c2c60dc9ba Removes multiple newlines at the end of files that is breaking the editorconfig step of CI. (llama/8258) Clint Herron 2024-07-02 12:18:10 -04:00
  • 4af3194b7c cuda : update supports_op for matrix multiplication (llama/8245) slaren 2024-07-02 08:39:38 +02:00
  • 4a2ba1a065 Fix win build conflict of math library (llama/8230) luoyu-intel 2024-07-02 04:50:07 +00:00
  • f096cc6807 Fix the sub group size of Intel (llama/8106) luoyu-intel 2024-07-02 02:16:00 +00:00
  • e4bc83ab47 CUDA: refactor and optimize IQ MMVQ (llama/8215) Johannes Gäßler 2024-07-01 20:39:06 +02:00
  • db7e0dbe6e Update SYCL-Rope op and Refactor (llama/8157) zhentaoyu 2024-07-01 19:39:06 +08:00
  • bf88c94da9 CUDA: fix MMQ stream-k for --split-mode row (llama/8167) Johannes Gäßler 2024-06-27 16:26:05 +02:00
  • 3eea171cab feat: cuda implementation for ggml_conv_transpose_1d (ggml/854) John Balis 2024-07-02 11:09:52 -05:00
  • 64a56ebf13 ci : disable java build Georgi Gerganov 2024-07-08 14:26:59 +03:00
  • bec9836849 server : add inference path to make OAI API compatible (#2270) Emmanuel Schmidbauer 2024-07-08 07:24:58 -04:00
  • c118733a29 sync : ggml + fix sync script Georgi Gerganov 2024-06-26 23:20:19 +03:00
  • bb3dd45524 make : disable CUDA graphs Georgi Gerganov 2024-06-26 23:20:13 +03:00
  • 04e7fa6f4f ggml : add GGML_CUDA_USE_GRAPHS option, restore GGML_CUDA_FORCE_CUBLAS (cmake) (llama/8140) slaren 2024-06-26 21:34:14 +02:00
  • 9f7f36d4c9 make : disable CUDA mel build Georgi Gerganov 2024-06-26 22:25:25 +03:00
  • 4a62efbb95 cmake : minor fixes Georgi Gerganov 2024-06-26 21:42:39 +03:00
  • 0a55a70b9b make : fix missing -O3 Georgi Gerganov 2024-06-26 21:20:45 +03:00
  • dc8cc2dd6f whisper : disable CUDA mel + fix FFMPEG Georgi Gerganov 2024-06-26 20:11:38 +03:00
  • 3efedb9511 sync : ggml Georgi Gerganov 2024-06-26 19:40:23 +03:00
  • e30c679928 whisper : reorganize source code + improve CMake (#2256) Georgi Gerganov 2024-06-26 19:34:09 +03:00
  • bf4cb4abad whisper : optimize fft() function (#2242) mky_coder 2024-06-18 23:10:33 +08:00
  • e293f17d34 talk-llama : sync llama.cpp Georgi Gerganov 2024-06-18 09:45:37 +03:00
  • 5d950c4b8d whisper : use ggml_backend_sched (#2239) Georgi Gerganov 2024-06-18 09:37:20 +03:00
  • 820446e230 fix : remove extra files Georgi Gerganov 2024-06-16 19:23:55 +03:00
  • 54d5823ebe scripts : sync ggml-blas Georgi Gerganov 2024-06-16 19:23:32 +03:00
  • 5181494e9f build : update make / cmake Georgi Gerganov 2024-06-16 19:10:20 +03:00
  • 4a6e6e8b30 sync : ggml Georgi Gerganov 2024-06-16 18:40:07 +03:00
  • de29b193f6 move BLAS to a separate backend (cont) (llama/6210) slaren 2024-06-16 13:57:37 +03:00
  • 922971041b Vulkan Shader Refactor, Memory Debugging Option (llama/7947) 0cc4m 2024-06-16 07:17:31 +02:00
  • 63a767a134 scripts : stop sync whisper example from ggml Georgi Gerganov 2024-06-16 18:38:46 +03:00
  • 30841fa786 cmake : fix sycl build (#0) Georgi Gerganov 2024-06-16 17:57:35 +03:00
  • 3b1ac03828 ggml : remove OpenCL (#0) Georgi Gerganov 2024-06-16 13:46:12 +03:00
  • 990de617b5 sycl : sync (#0) Georgi Gerganov 2024-06-16 13:24:17 +03:00
  • 6975600b4b cuda : enable CUDA graphs (#0) Georgi Gerganov 2024-06-16 13:20:19 +03:00
  • 061eeb9f61 talk-llama : sync llama.cpp Georgi Gerganov 2024-06-16 13:10:54 +03:00
  • 4942b1b428 cmake : fix CUDA build (#0) Georgi Gerganov 2024-06-16 13:07:43 +03:00
  • 3c7cc5c437 sync : ggml Georgi Gerganov 2024-06-16 12:43:14 +03:00
  • 5cd42ee2cc ggml : fix and optimize ppc64le (ggml/849) Hong Bo PENG 2024-06-16 16:53:11 +08:00
  • ee718f3da6 ggml : remove duplicate include of ggml-common.h (ggml/853) Daniel Bevenius 2024-06-16 10:51:18 +02:00
  • 63eac1f608 remove global variables (llama/7710) Meng, Hengyu 2024-06-15 14:05:10 +08:00
  • b17ba2815b CUDA: faster q2_K, q3_K MMQ + int8 tensor cores (llama/7921) Johannes Gäßler 2024-06-14 18:41:49 +02:00
  • 7a489af2f3 metal : utilize max shared memory for mul_mat_id (llama/7935) Georgi Gerganov 2024-06-14 17:14:09 +03:00
  • 4a4ea13d6d rpc : fix ggml_backend_rpc_supports_buft() (llama/7918) Radoslav Gerganov 2024-06-13 15:18:44 +03:00
  • 174a461fc6 move BLAS to a separate backend (llama/6210) slaren 2024-06-13 03:11:35 +02:00
  • d8b7a24bc9 CUDA: fix broken oob check for FA vec f32 kernel (llama/7904) Johannes Gäßler 2024-06-12 17:41:51 +02:00
  • acf3832c9c tests : add non-cont unary tests (llama/7857) Georgi Gerganov 2024-06-12 16:00:22 +03:00
  • d29ac44303 ggml : improve ggml_is_contiguous logic (llama/7856) Georgi Gerganov 2024-06-12 15:24:20 +03:00
  • 12638dfef0 vulkan: select only one device for single gpu with multiple drivers (llama/7582) k.h.lai 2024-06-12 03:26:05 +08:00
  • f100b3b523 Update Vulkan RoPE implementation (llama/7818) 0cc4m 2024-06-11 21:20:29 +02:00
  • a99e213a82 CUDA: int8 tensor cores for MMQ (q4_K, q5_K, q6_K) (llama/7860) Johannes Gäßler 2024-06-11 08:26:07 +02:00
  • 7483d2b61c CUDA: use tensor cores for MMQ (llama/7676) Johannes Gäßler 2024-06-10 11:45:13 +02:00
  • 1fe5948227 use the correct SYCL context for host USM allocations (llama/7777) Ben Ashbaugh 2024-06-10 02:21:31 -07:00
  • 760497e1ab CUDA: revise q8_1 data layout for mul_mat_q (llama/7824) Johannes Gäßler 2024-06-09 09:42:25 +02:00
  • b172e7714c vulkan : reuse parent extra for views (llama/7806) slaren 2024-06-07 19:47:49 +02:00
  • dc01aadb18 fix softmax r2r result wrong issue (llama/7811) pengxin99 2024-06-07 14:28:26 +08:00
  • e08c62149b CUDA: refactor mmq, dmmv, mmvq (llama/7716) Johannes Gäßler 2024-06-05 16:53:00 +02:00
  • abab4500fa ggml : refactor rope norm/neox (llama/7634) Georgi Gerganov 2024-06-05 11:29:20 +03:00
  • e666315fa8 Allow number of nodes in CUDA graph to change (llama/7738) agray3 2024-06-04 21:06:49 +01:00
  • 3f869af14c ggml : remove OpenCL (llama/7735) Georgi Gerganov 2024-06-04 21:23:20 +03:00
  • cbacb7634c ggml : prevent builds with -ffinite-math-only (llama/7726) Georgi Gerganov 2024-06-04 10:01:09 +03:00
  • 6cc3b022ee llama : offload to RPC in addition to other backends (llama/7640) Radoslav Gerganov 2024-06-03 20:03:26 +03:00
  • e5e38d4920 ggml : use OpenMP as a thread pool (llama/7606) Masaya, Kato 2024-06-04 00:14:15 +09:00
  • 2a6bab5655 Vulkan Mixture of Experts (MoE) support (llama/7628) 0cc4m 2024-06-03 10:59:14 +02:00
  • 8c01c9b85c kompute : implement op_getrows_f32 (llama/6403) woachk 2024-06-03 07:32:16 +02:00
  • d1123d795e fix bug introduced in using calloc (llama/7701) Dave Airlie 2024-06-03 07:59:54 +10:00
  • 9b3d784020 Fix FlashAttention debug test, FP32 assert (llama/7684) Johannes Gäßler 2024-06-01 23:26:10 +02:00
  • a16137d13d CUDA: fix Pascal FA, deq. KV to FP16 for batch > 8 (llama/7681) Johannes Gäßler 2024-06-01 15:47:04 +02:00
  • 5582039d0a CUDA: quantized KV support for FA vec (llama/7527) Johannes Gäßler 2024-06-01 08:44:14 +02:00
  • 9a16c643e2 ggml : fix loongson compile warnings (llama/7537) Georgi Gerganov 2024-05-31 14:17:10 +03:00
  • 10a8a23100 faster avx512 exp implementation (llama/7551) Chris Elrod 2024-05-30 07:32:55 -04:00
  • 29cfeef77f ggml : fix loongarch build (O2 issue) (llama/7636) junchao-loongson 2024-05-30 17:30:10 +08:00
  • e66e9ea25b metal : remove invalid asserts (llama/7617) Georgi Gerganov 2024-05-29 22:20:40 +03:00
  • 276779a849 metal : add missing asserts (llama/7617) Georgi Gerganov 2024-05-29 20:45:25 +03:00
  • 1f35ce61c1 ggml : fix YARN + add tests + add asserts (llama/7617) Georgi Gerganov 2024-05-29 20:17:31 +03:00