Commit Graph

  • 208de95ac7 conext add name (llama/5624) Meng, Hengyu 2024-02-21 17:52:06 +08:00
  • c2ce39c795 Update ggml_sycl_op_mul_mat_vec_q (llama/5502) AidanBeltonS 2024-02-20 07:01:25 +00:00
  • 8daa534818 Refactor validation and enumeration platform checks into functions to clean up ggml_vk_instance_init() 0cc4m 2024-02-14 20:57:17 +01:00
  • 9fca69b410 Add check for VK_KHR_portability_enumeration for MoltenVK support 0cc4m 2024-02-10 22:14:52 +01:00
  • b26c645420 Add preprocessor checks for Apple devices. Mathijs de Bruin 2024-02-06 14:39:22 +00:00
  • 1879ec556e Resolve ErrorIncompatibleDriver with Vulkan on MacOS. Mathijs de Bruin 2024-02-03 18:00:11 +00:00
  • c6e53cfc46 Allow for Vulkan build with Accelerate. Mathijs de Bruin 2024-02-03 17:56:46 +00:00
  • b19f2fb815 cuda : ignore peer access already enabled errors (llama/5597) slaren 2024-02-19 23:40:26 +01:00
  • a6b0950916 ggml : compute forward no longer pass src tensors (ggml/729) Siddharth Ramakrishnan 2024-02-21 04:34:53 -08:00
  • d352dbd163 ggml : fix conv_2d batch mode (ggml/737) bssrdf 2024-02-20 14:17:09 -05:00
  • eb23f4ef16 openvino : fix convert-whisper-to-openvino.py (#1890) st-gr 2024-02-22 05:11:35 -08:00
  • c56344b509 main : fix file existence check in main.cpp (#1889) Davidson Francis 2024-02-22 10:01:08 -03:00
  • 59119f4f20 talk-llama : sync llama.cpp Georgi Gerganov 2024-02-20 12:09:57 +02:00
  • 276615d708 make : fix CUBLAS link with WSL (#1878) LBlue 2024-02-20 18:05:38 +08:00
  • b602819b6e sync : ggml Georgi Gerganov 2024-02-19 15:54:25 +02:00
  • c2c606f05b ggml : resolve merge conflicts (ggml/0) Georgi Gerganov 2024-02-19 15:33:51 +02:00
  • 83afebe872 common : add IQ1_S (ggml/0) Georgi Gerganov 2024-02-19 15:27:37 +02:00
  • a4d8f9d559 ci : enable -Werror for CUDA builds (llama/5579) Georgi Gerganov 2024-02-19 14:45:41 +02:00
  • 5ec1e0edfa cuda, metal : fix nans in soft_max (llama/5574) slaren 2024-02-19 09:04:45 +01:00
  • 30a11b1ab8 ggml : android and old glibc NUMA incompatibility bugfixes (llama/5557) bmwl 2024-02-18 23:38:32 -08:00
  • f04e6b87d7 ggml : restore vec dot stride arg names (llama/5453) Georgi Gerganov 2024-02-18 22:58:57 +02:00
  • 0c33928b55 ci : fix wikitext url + compile warnings (llama/5569) Georgi Gerganov 2024-02-18 22:39:30 +02:00
  • 0775374750 metal : fix unused warnings (llama/0) Georgi Gerganov 2024-02-18 21:39:58 +02:00
  • 7d90bb035b ggml, common, examples, tests : fixed type arguments in printf (llama/5528) Herman Semenov 2024-02-18 16:20:12 +00:00
  • 2c1ad21ba8 1.5 bit quantization (llama/5453) Kawrakow 2024-02-18 18:16:55 +02:00
  • eca5ff9868 ggml : add ALiBi support for ggml_soft_max_ext (llama/5488) Georgi Gerganov 2024-02-19 15:18:09 +02:00
  • 1b25d2fa0a ci : add an option to fail on compile warning (llama/3952) Ananta Bastola 2024-02-17 16:03:14 -05:00
  • 74a6acc999 cmake : fix VULKAN and ROCm builds (llama/5525) Georgi Gerganov 2024-02-16 19:05:56 +02:00
  • a4ed8a0821 ggml : add numa options (llama/5377) bmwl 2024-02-16 01:31:07 -08:00
  • 9f675e021c cuda : print message when initialization fails (llama/5512) slaren 2024-02-15 16:49:01 +01:00
  • a38efcb9fd vulkan: Find optimal memory type but with fallback (llama/5381) Neuman Vong 2024-02-15 17:11:15 +11:00
  • 31591649a0 Early return for zero size calls to get_tensor. (llama/5482) AT 2024-02-13 15:44:25 -06:00
  • 4f5c46a84f ggml-quants : fix compiler warnings (shadow variable) (llama/5472) Kawrakow 2024-02-13 09:07:57 +02:00
  • 462ffc58db ggml-sycl: Replace 3d ops with macro (llama/5458) Abhilash Majumder 2024-02-12 20:22:05 +05:30
  • 65faae0b6a build : update CBLAS flags + fix unused var warning (#0) Georgi Gerganov 2024-02-19 14:44:46 +02:00
  • dda4b0ed06 main : check if input files exist before proceeding (#1872) Davidson Francis 2024-02-19 05:51:26 -03:00
  • 07d04280be examples : clean up common code (#1871) Felix 2024-02-19 09:50:15 +01:00
  • 917c56ded4 models : fix openvino setup info (#1874) Jumper775 2024-02-18 21:19:47 -05:00
  • 3d42463845 models : add update py requirements Georgi Gerganov 2024-02-13 11:51:32 +02:00
  • 3ffc83d90a swift : package no longer use ggml dependency (#1861) Georgi Gerganov 2024-02-12 19:54:11 +02:00
  • e3c5e2cba8 whisper : fix external encoder (#1860) Georgi Gerganov 2024-02-12 19:53:51 +02:00
  • b742f13e70 sync : ggml Georgi Gerganov 2024-02-12 19:07:56 +02:00
  • 52c529eeb1 ggml-alloc : allocate all leafs as if they were inputs (ggml/731) slaren 2024-02-12 18:07:14 +01:00
  • 551529290d talk-llama : sync llama.cpp Georgi Gerganov 2024-02-12 10:39:58 +02:00
  • 25a90ffa38 sync : ggml Georgi Gerganov 2024-02-12 09:32:15 +02:00
  • 866b67ca93 ggml-backend : sync remnant Georgi Gerganov 2024-02-12 09:27:57 +02:00
  • d7e9f58f7f CUDA: mul_mat_vec_q tiling, refactor mul mat logic (llama/5434) Johannes Gäßler 2024-02-11 19:08:39 +01:00
  • 04839bae22 vulkan: only use M-sized matmul on Apple GPUs (llama/5412) Sergio López 2024-02-11 15:12:00 +01:00
  • 3cc6e04a52 ggml : fix compile warnings (unused vars) (llama/4966) Georgi Gerganov 2024-02-11 15:33:01 +02:00
  • b7ef178b9c ggml : add mmla kernels for quantized GEMM (llama/4966) snadampal 2024-02-11 07:22:33 -06:00
  • 47dfe9d4db metal : use autoreleasepool to avoid memory leaks (llama/5437) Ian Bull 2024-02-10 02:53:28 -08:00
  • 1d3270cc8f ggml-alloc : v3 (ggml/727) slaren 2024-02-11 13:37:58 +01:00
  • a6fb6ab597 examples : added audio_ctx argument to main and server (#1857) dscripka 2024-02-12 02:19:07 -05:00
  • 163e74b6c3 metal : option to embed MSL source into compiled binary (#1842) Didzis Gosko 2024-02-11 16:41:41 +02:00
  • f273e66dc6 examples : initialize context params properly (#1852) Georgi Gerganov 2024-02-11 16:39:12 +02:00
  • 02b4c52c12 talk-llama : sync llama.cpp Georgi Gerganov 2024-02-10 10:10:59 +02:00
  • 518199c09e sync : ggml Georgi Gerganov 2024-02-10 09:56:47 +02:00
  • 8b17a2f776 src : relocate new backend sources Georgi Gerganov 2024-02-10 09:50:24 +02:00
  • b6d2827914 ggml : fix error C2078: too many initializers for MSVC ARM64 (llama/5404) Michael Podvitskiy 2024-02-09 10:56:43 +01:00
  • 9711bae0b3 CUDA: more warps for mmvq on NVIDIA (llama/5394) Johannes Gäßler 2024-02-08 21:56:40 +01:00
  • eec38f63bd CUDA: fixed mmvq kernel for bs 2,3,4 and -sm row (llama/5386) Johannes Gäßler 2024-02-07 12:40:26 +01:00
  • ef5e6b746f Basic Vulkan Multi-GPU implementation (llama/5321) 0cc4m 2024-02-07 07:54:50 +01:00
  • 77bf6b5f56 CUDA: mul_mat_vec_q max. batch size 8 -> 4 (llama/5370) Johannes Gäßler 2024-02-06 18:43:06 +01:00
  • b562fff9d0 Slight quantization improvement for Q4_K and Q5_K (llama/5361) Kawrakow 2024-02-06 17:28:02 +02:00
  • b5dec374f4 CUDA: mul_mat_vec_q for batch sizes > 1 (llama/5351) Johannes Gäßler 2024-02-06 14:44:06 +01:00
  • fa0dc6167c ggml : make use of ggml-quants.h possible in C++ code (llama/5338) Kawrakow 2024-02-05 14:09:47 +02:00
  • 55bcd62a4b ggml : avoid duplicating function calls using MIN/MAX macros (llama/5325) Dr. Tom Murphy VII Ph.D 2024-02-05 06:13:57 -05:00
  • 0ed762d691 iq2_xxs: tune quantization (llama/5320) Kawrakow 2024-02-05 10:46:06 +02:00
  • 1b5bb7792e cuda : fix LLAMA_CUDA_F16 (llama/5262) slaren 2024-02-01 18:30:17 +01:00
  • 9b735cea77 metal : add im2col F32 dst support (llama/5132) Georgi Gerganov 2024-01-31 15:35:41 +02:00
  • 12c462d656 llava : add MobileVLM support (llama/5132) JidongZhang-THU 2024-01-31 21:10:15 +08:00
  • fc7b0e2c28 ggml : limit n_threads to the max n_tasks (llama/5238) slaren 2024-01-31 13:43:03 +01:00
  • f850a067ed kompute : llama-bench support and ggml_cpu_has_kompute() (llama/5226) Jared Van Bortel 2024-01-30 19:04:37 -05:00
  • f75e1197f1 ggml : add abort_callback for cpu backend (ggml/725) Michael Podvitskiy 2024-02-09 10:42:27 +01:00
  • aa8a75e287 extra : update sync scripts Georgi Gerganov 2024-02-10 09:55:19 +02:00
  • 80e8a2ea39 server : allow CORS request with authorization headers (#1850) Valentin Gosu 2024-02-09 16:42:41 +01:00
  • 19f8048139 whisper.android : how to build with CLBlast (#1809) Neuman Vong 2024-02-10 02:39:05 +11:00
  • 0f80e5a80a whisper : expose CUDA device setting in public API (#1840) Didzis Gosko 2024-02-09 17:27:47 +02:00
  • b6559333ff make : add macOS deployment target option (#1839) Didzis Gosko 2024-02-09 17:26:29 +02:00
  • 434b8f3b96 talk-llama : stream response (#1121) Georgi Gerganov 2024-02-06 19:56:12 +02:00
  • 7a74e929c8 sync : ggml (#0) Georgi Gerganov 2024-01-30 21:30:26 +02:00
  • 361ecebe90 ggml : fix IQ3_XXS on Metal (llama/5219) Kawrakow 2024-01-30 19:15:28 +02:00
  • 807cbc672e sync : ggml (llama/0) Georgi Gerganov 2024-01-30 16:21:57 +02:00
  • 98ae5276b7 Faster AVX2 dot product for IQ2_XS (llama/5187) Kawrakow 2024-01-30 15:15:07 +02:00
  • 6adb969b09 SOTA 3-bit quants (llama/5196) Kawrakow 2024-01-30 15:14:12 +02:00
  • 8a7d6ff51a ggml alloc: Fix for null dereference on alloc failure (llama/5200) Paul Tsochantaris 2024-01-29 22:19:29 +00:00
  • 25f650a8e8 Nomic Vulkan backend (llama/4456) Jared Van Bortel 2024-01-29 15:50:50 -05:00
  • 44e517f074 ggml : add max buffer sizes to opencl and metal backends (llama/5181) slaren 2024-01-29 09:05:13 +01:00
  • cb9de61659 metal : free metal objects (llama/5161) Paul Tsochantaris 2024-01-28 19:50:16 +00:00
  • a2ef80d66f gguf : fix comparison (ggml/715) Georgi Gerganov 2024-01-29 21:08:18 +02:00
  • baa190446a ggml_cuda_cpy support for 4d tensors and float16->float32 upcasting (ggml/686) John Balis 2024-01-29 06:37:33 -06:00
  • 8f5220d81f gguf : add input validation, prevent integer overflows (ggml/709) Georgi Gerganov 2024-01-29 14:00:10 +02:00
  • 8e391fcf3a ci : fix yolo URLs + fix metal capture (ggml/712) Georgi Gerganov 2024-01-29 13:29:46 +02:00
  • 593657054e metal : add debug capture backend function (ggml/694) Jack Mousseau 2024-01-29 01:22:23 -08:00
  • ae5c4f7340 common : fix wav buffer detection (#1819) JacobLinCool 2024-01-31 01:35:08 +08:00
  • baa30bacdb server : add fields to verbose_json response (#1802) JacobLinCool 2024-01-30 20:15:55 +08:00
  • 3e6fad07aa make : update MSYS_NT (#1813) jwijffels 2024-01-30 13:13:49 +01:00
  • e72e4158de talk-llama : sync llama.cpp Georgi Gerganov 2024-01-28 19:44:10 +02:00
  • bd41733db2 sync : ggml Georgi Gerganov 2024-01-28 19:30:32 +02:00
  • 23c648e98d ggml : add Vulkan backend (llama/2059) 0cc4m 2024-01-28 18:03:59 +01:00