leejet
1330cebae8
feat: support building with upstream ggml ( #1999 )
2026-09-19 22:25:01 +08:00
leejet
17860c0e45
perf: parallelize host tensor elementwise and broadcast ops ( #1998 )
2026-09-19 21:46:16 +08:00
leejet
9982c9caae
fix: propagate CUDA driver dependency to shared library consumers
2026-09-19 18:04:59 +08:00
leejet
adcac69650
perf: pad small attention heads to 64 for MMA Flash Attention ( #1992 )
2026-09-18 23:59:58 +08:00
leejet
4964abdfc5
feat: support external Hugging Face tokenizer JSON files ( #1973 )
2026-09-15 00:27:39 +08:00
leejet
9a977388a8
fix: guard GPU memory capacity and propagate encoding failures ( #1958 )
2026-09-13 21:35:46 +08:00
leejet
14eddb32b1
refactor: split generation pipeline out of stable-diffusion.cpp ( #1957 )
2026-09-11 00:57:14 +08:00
assouan
6c57cc3b38
feat: prefetch streamed layers during compute ( #1905 )
...
Co-authored-by: leejet <leejet714@gmail.com>
2026-09-06 16:35:30 +08:00
Huang, Hong-Chang
b4e67d1221
fix(cmake): only apply /MP to the MSVC compiler, not icx ( #1846 )
2026-08-04 22:37:30 +08:00
fszontagh
c00a9e956d
feat: AnimateDiff SD 1.5 motion modules (v2 + v3) ( #1784 )
2026-07-14 23:58:03 +08:00
Cyberhan123
c1790754d3
feat: enhanced third-party integrations ( #1632 )
...
* feat: add installation support and configuration files for stable-diffusion
* fix: correct public header setting and update version variable in pkg-config
* fix stable-diffusion install package metadata
---------
Co-authored-by: leejet <leejet714@gmail.com>
2026-06-29 00:48:57 +08:00
stduhpf
c2df4e1228
feat: add RPC support ( #1629 )
2026-06-14 17:30:23 +08:00
leejet
2a07540c2a
refactor: move photomaker into generation extension ( #1618 )
2026-06-07 22:40:02 +08:00
Wagner Bruna
81abfb2548
chore: rename and reformat gits_noise.inl ( #1617 )
2026-06-07 22:30:20 +08:00
leejet
f3fd359b58
refactor: reorganize src model layout ( #1615 )
2026-06-07 03:21:12 +08:00
RapidMark
7948df8ac1
fix(cmake): build HIP backend with PIC so the static-lib PIE link succeeds ( #1593 )
2026-06-02 00:07:48 +08:00
leejet
72e512a0cc
fix: make macOS binaries use relocatable rpaths ( #1552 )
2026-05-23 12:27:06 +08:00
leejet
67dda3f897
feat: add ltx2.3 support ( #1463 )
...
* add GemmaTokenizer
* add basic ltx2.3 support
* change vocab file encoding
* fix ci
* fix ubuntu build
* add temporal tiling support
* add ltx audio support
* update ggml submodule url
* fix generate_video
* add i2v support
* minify bundled Gemma tokenizer vocab sources
* pass video fps into temporal rope embeddings
* fix av_ca_timestep_scale_multiplier
* add LTX2Scheduler support
* update docs
* fix ci
2026-05-17 16:46:20 +08:00
cphlipot
db08b84607
fix: Fix broken GCC 16 build (enforce C11/C++17 compile ) ( #1478 )
2026-05-16 16:10:16 +08:00
Craig Andrews
eeac950b44
fix: Use PkgConfig for WebP and WebM ( #1400 )
2026-05-15 00:31:10 +08:00
Wagner Bruna
b8079e253d
feat: transition from compile-time to runtime backend discovery ( #1448 )
...
Co-authored-by: Stéphane du Hamel <stephduh@live.fr>
Co-authored-by: Cyberhan123 <255542417@qq.com>
Co-authored-by: leejet <leejet714@gmail.com>
2026-04-29 23:26:57 +08:00
leejet
66143340b6
refactor: move model file IO into dedicated module ( #1442 )
2026-04-19 17:52:56 +08:00
leejet
7d33d4b2dd
chore: enable MSVC parallel compilation with /MP ( #1438 )
2026-04-18 15:44:43 +08:00
leejet
9ac7b672c2
refactor: introduce shared tokenizer abstraction and split implementations ( #1423 )
2026-04-15 22:44:39 +08:00
leejet
7397ddaa86
feat: add webm support ( #1391 )
2026-04-06 01:49:28 +08:00
Wagner Bruna
687a81f251
chore: make libwebp optional and support system libwebp ( #1387 )
...
Co-authored-by: leejet <leejet714@gmail.com>
2026-04-05 23:52:05 +08:00
leejet
87ecb95cbc
feat: add webp support ( #1384 )
2026-04-02 01:36:11 +08:00
Wagner Bruna
f6968bc589
chore: remove SD_FAST_SOFTMAX build flag ( #1338 )
2026-03-15 16:42:47 +08:00
leejet
636d3cb6ff
refactor: reorganize the vocab file structure ( #1271 )
2026-02-11 00:44:17 +08:00
leejet
28ef93c0e1
refactor: reorganize the file structure ( #1266 )
2026-02-10 23:13:35 +08:00
leejet
b90b1ee9cf
chore: eliminate compilation warnings under MSVC ( #1170 )
2026-01-04 22:26:57 +08:00
Wagner Bruna
e72aea796e
feat: embed version string and git commit hash ( #1008 )
2025-12-09 22:38:54 +08:00
cmdr2
a7d6d296c7
chore: allow building ggml as a separate shared lib ( #468 )
2025-10-15 22:10:26 +08:00
clibdev
6bbaf161ad
chore: add install() support in CMakeLists.txt ( #540 )
2025-09-11 22:24:16 +08:00
leejet
675208dcb6
chore: update to c++17
2025-09-07 12:04:17 +08:00
Wagner Bruna
f7f05fb185
chore: avoid setting GGML_MAX_NAME when building against external ggml ( #751 )
...
An external ggml will most likely have been built with the default
GGML_MAX_NAME value (64), which would be inconsistent with the value
set by our build (128). That would be an ODR violation, and it could
easily cause memory corruption issues due to the different
sizeof(struct ggml_tensor) values.
For now, when linking against an external ggml, we demand it has been
patched with a bigger GGML_MAX_NAME, since we can't check against a
value defined only at build time.
2025-08-03 01:24:40 +08:00
Seas0
6167e2927a
feat: support build against system installed GGML library ( #749 )
2025-08-02 11:03:18 +08:00
rmatif
d42fd59464
feat: add OpenCL backend support ( #680 )
2025-06-30 23:32:23 +08:00
Meng, Hengyu
838beb9b5e
chore: add global SYCL compile flags ( #597 )
2025-02-22 21:23:58 +08:00
R0CKSTAR
a3cbdf6dcb
chore: SD_USE_CUBLAS => SD_USE_CUDA for MUSA backend ( #578 )
...
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2025-02-05 16:11:26 +08:00
null-define
b70aaa672a
chore: fix amd rocm build ( #571 )
2025-01-18 13:11:39 +08:00
leejet
dcf91f9e0f
chore: change SD_CUBLAS/SD_USE_CUBLAS to SD_CUDA/SD_USE_CUDA
2024-12-28 13:27:51 +08:00
R0CKSTAR
5cc74d1f09
feat: support Moore Threads GPU ( #529 )
...
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
2024-12-28 13:08:36 +08:00
Erik Scholz
1c168d98a5
fix: repair flash attention support ( #386 )
...
* repair flash attention in _ext
this does not fix the currently broken fa behind the define, which is only used by VAE
Co-authored-by: FSSRepo <FSSRepo@users.noreply.github.com>
* make flash attention in the diffusion model a runtime flag
no support for sd3 or video
* remove old flash attention option and switch vae over to attn_ext
* update docs
* format code
---------
Co-authored-by: FSSRepo <FSSRepo@users.noreply.github.com>
Co-authored-by: leejet <leejet714@gmail.com>
2024-11-23 12:39:08 +08:00
zhentaoyu
e410aeb534
sync: update ggml to fix large image generation with SYCL backend ( #380 )
...
* turn off fast-math on host in SYCL backend
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* update ggml for sync some sycl ops
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* update sycl readme and ggml
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
---------
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
2024-09-02 22:29:35 +08:00
Yu Xing
6c88ad3fd6
fix: resolve naming conflict while llama.cpp and sd.cpp both build ( #351 )
2024-08-28 00:14:41 +08:00
soham
2027b16fda
feat: add vulkan backend support ( #291 )
...
* Fix includes and init vulkan the same as llama.cpp
* Add Windows Vulkan CI
* Updated ggml submodule
* support epsilon as a parameter for ggml_group_norm
---------
Co-authored-by: Cloudwalk <cloudwalk@icculus.org>
Co-authored-by: Oleg Skutte <00.00.oleg.00.00@gmail.com>
Co-authored-by: leejet <leejet714@gmail.com>
2024-08-27 23:56:09 +08:00
zhentaoyu
697d000f49
feat: add SYCL Backend Support for Intel GPUs ( #330 )
...
* update ggml and add SYCL CMake option
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* hacky CMakeLists.txt for updating ggml in cpu backend
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* rebase and clean code
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* add sycl in README
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* rebase ggml commit
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* refine README
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
* update ggml for supporting sycl tsembd op
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
---------
Signed-off-by: zhentaoyu <zhentao.yu@intel.com>
2024-08-10 13:42:50 +08:00
leejet
be6cd1a4bf
sync: update ggml
2024-06-01 13:44:09 +08:00
Phu Tran
1ce9470f27
fix: fix building shared library ( #188 )
2024-03-03 13:24:59 +08:00