rocm-stable-diffusion.cpp

mirror of https://github.com/BillyOutlast/rocm-stable-diffusion.cpp.git synced 2026-02-04 11:11:19 +01:00

Author	SHA1	Message	Date
Wagner Bruna	48956ffb87	feat: reduce CLIP memory usage with no embeddings (#768 )	2025-09-14 12:08:00 +08:00
leejet	fce6afcc6a	feat: add sd3 flash attn support (#815 )	2025-09-11 23:24:29 +08:00
Erik Scholz	49d6570c43	feat: add SmoothStep Scheduler (#813 )	2025-09-11 23:17:46 +08:00
clibdev	87cdbd5978	feat: use log_printf to print ggml logs (#545 )	2025-09-11 22:16:05 +08:00
leejet	b017918106	chore: remove sd3 flash attention warn (#812 )	2025-09-10 22:21:02 +08:00
Wagner Bruna	ac5a215998	fix: use {} for params init instead of memset (#781 )	2025-09-10 21:49:29 +08:00
Wagner Bruna	abb36d66b5	chore: update flash attention warnings (#805 )	2025-09-10 21:38:21 +08:00
Wagner Bruna	ff4fdbb88d	fix: accept NULL in sd_img_gen_params_t::input_id_images_path (#809 )	2025-09-10 21:22:55 +08:00
Markus Hartung	abb115cd02	fix: clarify lora quant support and small fixes (#792 )	2025-09-08 22:39:25 +08:00
leejet	c648001030	feat: add detailed tensor loading time stat (#793 )	2025-09-07 22:51:44 +08:00
stduhpf	c587a43c99	feat: support incrementing ref image index (omni-kontext) (#755 ) * kontext: support ref images indices * lora: support x_embedder * update help message * Support for negative indices * support for OmniControl (offsets at index 0) * c++11 compat * add --increase-ref-index option * simplify the logic and fix some issues * update README.md * remove unused variable --------- Co-authored-by: leejet <leejet714@gmail.com>	2025-09-07 22:35:16 +08:00
stduhpf	141a4b4113	feat: add flow shift parameter (for SD3 and Wan) (#780 ) * Add flow shift parameter (for SD3 and Wan) * unify code style and fix some issues --------- Co-authored-by: leejet <leejet714@gmail.com>	2025-09-07 02:16:59 +08:00
stduhpf	21ce9fe2cf	feat: add support for timestep boundary based automatic expert routing in Wan MoE (#779 ) * Wan MoE: Automatic expert routing based on timestep boundary * unify code style and fix some issues --------- Co-authored-by: leejet <leejet714@gmail.com>	2025-09-07 01:44:10 +08:00
leejet	cb1d975e96	feat: add wan2.1/2.2 support (#778 ) * add wan vae suppport * add wan model support * add umt5 support * add wan2.1 t2i support * make flash attn work with wan * make wan a little faster * add wan2.1 t2v support * add wan gguf support * add offload params to cpu support * add wan2.1 i2v support * crop image before resize * set default fps to 16 * add diff lora support * fix wan2.1 i2v * introduce sd_sample_params_t * add wan2.2 t2v support * add wan2.2 14B i2v support * add wan2.2 ti2v support * add high noise lora support * sync: update ggml submodule url * avoid build failure on linux * avoid build failure * update ggml * update ggml * fix sd_version_is_wan * update ggml, fix cpu im2col_3d * fix ggml_nn_attention_ext mask * add cache support to ggml runner * fix the issue of illegal memory access * unify image loading processing * add wan2.1/2.2 FLF2V support * fix end_image mask * update to latest ggml * add GGUFReader * update docs	2025-09-06 18:08:03 +08:00
Daniele	5b8996f74a	Conv2D direct support (#744 ) * Conv2DDirect for VAE stage * Enable only for Vulkan, reduced duplicated code * Cmake option to use conv2d direct * conv2d direct always on for opencl * conv direct as a flag * fix merge typo * Align conv2d behavior to flash attention's * fix readme * add conv2d direct for controlnet * add conv2d direct for esrgan * clean code, use enable_conv2d_direct/get_all_blocks * format code --------- Co-authored-by: leejet <leejet714@gmail.com>	2025-08-03 01:25:17 +08:00
stduhpf	59080d3ce1	feat: change image dimensions requirement for DiT models (#742 )	2025-07-28 21:58:17 +08:00
Oleg Skutte	1896b28ef2	fix: make --taesd work (#731 )	2025-07-15 00:45:22 +08:00
leejet	ca0bd9396e	refactor: update c api (#728 )	2025-07-13 18:48:42 +08:00
stduhpf	a772dca27a	feat: add Instruct-Pix2pix/CosXL-Edit support (#679 ) * Instruct-p2p support * support 2 conditionings cfg * Do not re-encode the exact same image twice * fixes for 2-cfg * Fix pix2pix latent inputs + improve inpainting a bit + fix naming * prepare for other pix2pix-like models * Support sdxl ip2p * fix reference image embeddings * Support 2-cond cfg properly in cli * fix typo in help * Support masks for ip2p models * unify code style * delete unused code * use edit mode * add img_cond * format code --------- Co-authored-by: leejet <leejet714@gmail.com>	2025-07-12 15:36:45 +08:00
stduhpf	19fbfd8639	feat: override text encoders for unet models (#682 )	2025-07-04 22:19:47 +08:00
leejet	7dac89ad75	refector: reuse some code	2025-07-01 23:33:50 +08:00
stduhpf	9251756086	feat: add CosXL support (#683 )	2025-07-01 23:13:04 +08:00
leejet	23de7fc44a	chore: avoid warnings when building on linux	2025-06-30 23:49:52 +08:00
rmatif	d42fd59464	feat: add OpenCL backend support (#680 )	2025-06-30 23:32:23 +08:00
leejet	45d0ebb30c	style: format code	2025-06-29 23:40:55 +08:00
stduhpf	b1cc40c35c	feat: add Chroma support (#696 ) --------- Co-authored-by: Green Sky <Green-Sky@users.noreply.github.com> Co-authored-by: leejet <leejet714@gmail.com>	2025-06-29 23:36:42 +08:00
stduhpf	c9b5735116	feat: add FLUX.1 Kontext dev support (#707 ) * Kontext support * add edit mode --------- Co-authored-by: leejet <leejet714@gmail.com>	2025-06-29 10:08:53 +08:00
vmobilis	10feacf031	fix: correct img2img time (#616 )	2025-03-09 12:29:08 +08:00
yslai	19d876ee30	feat: implement DDIM with the "trailing" timestep spacing and TCD (#568 )	2025-02-22 21:34:22 +08:00
stduhpf	f23b803a6b	fix:: unapply current loras properly (#590 )	2025-02-22 21:22:22 +08:00
stduhpf	d9b5942d98	feat: add sdxl v-pred suppport (#536 )	2025-01-18 13:15:54 +08:00
stduhpf	587a37b2e2	fix: avoid sd2((non inpaint) crash on v-pred check (#537 )	2025-01-18 13:13:34 +08:00
leejet	dcf91f9e0f	chore: change SD_CUBLAS/SD_USE_CUBLAS to SD_CUDA/SD_USE_CUDA	2024-12-28 13:27:51 +08:00
stduhpf	d50473dc49	feat: support 16 channel tae (taesd/taef1) (#527 )	2024-12-28 13:13:48 +08:00
stduhpf	0d9d6659a7	fix: fix metal build (#513 )	2024-12-28 13:06:17 +08:00
stduhpf	8f4ab9add3	feat: support Inpaint models (#511 )	2024-12-28 13:04:49 +08:00
stduhpf	cc92a6a1b3	feat: support more LoRA models (#520 )	2024-12-28 12:56:44 +08:00
stduhpf	7ce63e740c	feat: flexible model architecture for dit models (Flux & SD3) (#490 ) * Refactor: wtype per tensor * Fix default args * refactor: fix flux * Refactor photmaker v2 support * unet: refactor the refactoring * Refactor: fix controlnet and tae * refactor: upscaler * Refactor: fix runtime type override * upscaler: use fp16 again * Refactor: Flexible sd3 arch * Refactor: Flexible Flux arch * format code --------- Co-authored-by: leejet <leejet714@gmail.com>	2024-11-30 14:18:53 +08:00
stduhpf	53b415f787	fix: remove default variables in c headers (#478 )	2024-11-24 18:10:25 +08:00
leejet	b5f4932696	refactor: add some sd vesion helper functions	2024-11-23 13:02:44 +08:00
Erik Scholz	1c168d98a5	fix: repair flash attention support (#386 ) * repair flash attention in _ext this does not fix the currently broken fa behind the define, which is only used by VAE Co-authored-by: FSSRepo <FSSRepo@users.noreply.github.com> * make flash attention in the diffusion model a runtime flag no support for sd3 or video * remove old flash attention option and switch vae over to attn_ext * update docs * format code --------- Co-authored-by: FSSRepo <FSSRepo@users.noreply.github.com> Co-authored-by: leejet <leejet714@gmail.com>	2024-11-23 12:39:08 +08:00
bssrdf	2b1bc06477	feat: add PhotoMaker Version 2 support (#358 ) * first attempt at updating to photomaker v2 * continue adding photomaker v2 modules * finishing the last few pieces for photomaker v2; id_embeds need to be done by a manual step and pass as an input file * added a name converter for Photomaker V2; build ok * more debugging underway * failing at cuda mat_mul * updated chunk_half to be more efficient; redo feedforward * fixed a bug: carefully using ggml_view_4d to get chunks of a tensor; strides need to be recalculated or set properly; still failing at soft_max cuda op * redo weight calculation and weightv fixed a bug now Photomaker V2 kinds of working * add python script for face detection (Photomaker V2 needs) * updated readme for photomaker * fixed a bug causing PMV1 crashing; both V1 and V2 work * fixed clean_input_ids for PMV2 * fixed a double counting bug in tokenize_with_trigger_token * updated photomaker readme * removed some commented code * improved reconstructing class word free prompt * changed reading id_embed to raw binary using existing load tensor function; this is more efficient than using model load and also makes it easier to work with sd server * minor clean up --------- Co-authored-by: bssrdf <bssrdf@gmail.com>	2024-11-23 11:50:14 +08:00
stduhpf	6ea812256e	feat: add flux 1 lite 8B (freepik) support (#474 ) * Flux Lite (Freepik) support * format code --------- Co-authored-by: leejet <leejet714@gmail.com>	2024-11-23 11:41:30 +08:00
stduhpf	65fa646684	feat: add sd3.5 medium and skip layer guidance support (#451 ) * mmdit-x * add support for sd3.5 medium * add skip layer guidance support (mmdit only) * ignore slg if slg_scale is zero (optimization) * init out_skip once * slg support for flux (expermiental) * warn if version doesn't support slg * refactor slg cli args * set default slg_scale to 0 (oops) * format code --------- Co-authored-by: leejet <leejet714@gmail.com>	2024-11-23 11:15:31 +08:00
leejet	ac54e00760	feat: add sd3.5 support (#445 )	2024-10-24 21:58:03 +08:00
Daniele	0362cc4874	fix: fix some typos (#361 )	2024-08-28 00:15:37 +08:00
Daniele	dc0882cdc9	feat: add exponential scheduler (#346 ) * feat: added exponential scheduler * updated README * improved exponential formatting --------- Co-authored-by: leejet <leejet714@gmail.com>	2024-08-28 00:13:35 +08:00
Daniele	d00c94844d	feat: add ipndm and ipndm_v samplers (#344 )	2024-08-28 00:03:41 +08:00
Daniele	2d4a2f7982	feat: add GITS scheduler (#343 )	2024-08-28 00:02:17 +08:00
soham	2027b16fda	feat: add vulkan backend support (#291 ) * Fix includes and init vulkan the same as llama.cpp * Add Windows Vulkan CI * Updated ggml submodule * support epsilon as a parameter for ggml_group_norm --------- Co-authored-by: Cloudwalk <cloudwalk@icculus.org> Co-authored-by: Oleg Skutte <00.00.oleg.00.00@gmail.com> Co-authored-by: leejet <leejet714@gmail.com>	2024-08-27 23:56:09 +08:00

1 2 3 4

168 Commits