github/ROCm - ROCm - AtHeartEngineering

mirror of https://github.com/ROCm/ROCm.git synced 2026-04-05 03:01:17 -04:00

Author	SHA1	Message	Date
Philippe Tillet	20100a7254	Merge `triton-mlir` branch - Complete rewrite of the backend from scratch (#1004 ) This PR merges the `triton-mlir` branch, in which we have been quietly rewriting the Triton backend from scratch to increase maintainability, stability and ultimately performance. Changes to the runtime are minimal, and this new version aims to remain backward-compatible with the previous commit. The legacy backend is now officially deprecated, but can still be accessed via the `legacy-backend` tag. Co-authored-by: Keren Zhou <kerenzhou@openai.com> Co-authored-by: Yan Chunwei <yanchunwei@outlook.com> Co-authored-by: goostavz <109190422+goostavz@users.noreply.github.com> Co-authored-by: Shintaro Iwasaki <siwasaki@fb.com> Co-authored-by: Yan Da <dyanab@connect.ust.hk> Co-authored-by: Jun Yang <yangjunpro@gmail.com> Co-authored-by: Ian Bearman <ianb@microsoft.com> Co-authored-by: Jason Ansel <jansel@jansel.net> Co-authored-by: Qingyi Liu <qingyil@nvidia.com> Co-authored-by: ben-zhang-609 <110140741+ben-zhang-609@users.noreply.github.com> Co-authored-by: Chenggang Zhao <lyricz@yeah.net> Co-authored-by: ben-zhang-609 <benzh609@gmail.com> Co-authored-by: dongdongl <dongdongl@nvidia.com>	2022-12-21 01:30:50 -08:00
Yang Hau	8650b4d1cb	[DRIVER] Fix typos (#939 )	2022-12-02 11:13:46 -08:00
Yanbo Liang	5ca1ed0101	Add bf16/fp16/fp64 support for ty_to_cpp (#800 ) In ```torch._inductor```, we [convert 0d CPU tensor to scalar during triton codegen](https://github.com/pytorch/pytorch/pull/87329), so need add missing triton support for bf16/fp16/fp64.	2022-10-24 19:41:25 -07:00
Philippe Tillet	af76c989eb	[RUNTIME] Make entry point cache key depend on triton version hash (#765 )	2022-10-11 13:24:30 -07:00
Keren Zhou	11345e9b74	[RUNTIME] Add callback functions for external tools (#738 )	2022-10-05 14:46:55 -07:00
Philippe Tillet	bdfdb9a1d2	[RUNTIME] Fixed JIT bug that leg some constexpr values to be overriden by specialization parameters (#742 )	2022-10-05 11:00:32 -07:00
shenggan	77c752dc78	[RUNTIME] remove fixed cu_include_dir (#739 ) Use environment variable `CUDA_HOME` with default value`/usr/local/cuda` for `cu_include_dir` #731	2022-10-04 19:49:57 -07:00
fdrocha	2b0f877fad	[RUNTIME] Support environments with multiple cudalibs (#733 )	2022-10-03 18:36:24 +00:00
Keren Zhou	4a2d3b7d79	[RUNTIME] Dump llvm, ttir, and sass to help debugging (#732 )	2022-10-03 00:39:52 +00:00
Jason Ansel	998fd5f9af	[FRONTEND] Make triton.compile work without a cuda context (#708 ) This allows compiling in a subprocess. I'm not seeing a ton of speedup from this, but figure it is a good change anyway.	2022-09-24 13:41:47 -07:00
Philippe Tillet	8c3d4d5749	[RUNTIME] now decoupling entry point from cubin (#696 )	2022-09-22 16:44:22 -07:00
Jason Ansel	6abe813d1c	Fix issue breaking cudagraphs (#685 ) @ngimel figured this one out. The errors we were seeing from cudagraphs capture were coming from `cuStreamGetCtx` which is not allowed while a stream is capturing. It appears the result of `cuStreamGetCtx()` isn't even used, so I believe it can just be removed.	2022-09-21 10:20:48 -07:00
Philippe Tillet	48f30550f1	[FRONTEND] Now using raw compiler syscalls when possible (#678 )	2022-09-19 21:01:36 -07:00
Jason Ansel	49f6bc3f2b	[FRONTEND] Fix filename too long error in new runtime (#669 )	2022-09-18 21:26:29 +00:00
Jason Ansel	e647402fd3	Fix warning in generated C code (#667 )	2022-09-18 12:57:32 -07:00
Philippe Tillet	4a77dfb042	[FRONTEND] Complete rewrite of the runtime (#644 ) This PR completely rewrites the runtime of Triton to be more lean and clearly separate the compilation step from the just-in-time caching logic. This should substantially reduce launch overhead.	2022-09-18 08:51:48 -07:00

16 Commits