tinygrad

mirror of https://github.com/tinygrad/tinygrad.git synced 2026-02-19 02:44:40 -05:00

Author	SHA1	Message	Date
qazal	a5204fe89d	refactor UOps.CONST (#4639 ) * delete more * nit: dont need assign * can this be simpler * use scalars * always cast * clang needs cast * format	2024-05-18 10:07:36 +03:00
George Hotz	b74cc1d01a	uops cleanup (#4634 ) * def add cleanup * minor speedup * add back ptx speed * a little faster * merge that * only linearize once for ptx * two graph rewrites for ptx, bug?	2024-05-17 20:02:38 -07:00
George Hotz	07b350a8f4	new uops is an actual graph (#4560 ) * new uops is an actual graph * it's way slower * simpler * fix define acc * render_loop unique * ops test pass * add pattern matcher back, there's bugs * rewrite * use priority queue * recursive children * fix tests * fix tests with SINK * fix abstractions * fix assembly * simpler * link define_acc * fix DEFINE_ACC placement * type verify * full cmp * fix cmp * ACCESS_ACC * insert DEFINE_ACC * fix PHI * recursive rewrite * fix many tests * sum collapse * more patterns * correct change * fold arange * fix that lin test * space * big folding rule works * close * has more maxes, meh * cached node replace * set changed * simplest folding yet * works * works * DIV * all tests pass * del * fuzz linearizer fails * sum_collapse * test depth 2 cf * fix lin test 14 * fix clang depth * disable that * failure 14 is fixed * fix ptx * failure 27 is fixed * fix llama * run_cnt * Revert "Optimize PTX gated loads index calculation (#4304)" This reverts commit `d97d5a7689`. * fix uops loop * fix ptx bugs * add barrier * print * mem_type in ptx direct * bypass tests that fail in CI but pass locally * ptx remove ptr_ar * more ptx passing * fix ptx tests * assert compile support * remove model inference benchmark from red	2024-05-17 18:00:18 -07:00
nimlgen	daf57af3eb	move tc to renderers (#4631 ) * move tc to renderers * missed import * fix typo * fix * fix imports * remove from tests * fix 4607 * nv emulate timestamp * time is int * correct time	2024-05-18 00:36:29 +03:00
Szymon Ożóg	d97d5a7689	Optimize PTX gated loads index calculation (#4304 ) * WIP but working * Cleanup * Remove float4 pred and alt * Cleanup * this is somehow slowin it down * Simplify * add define var to ignore when optimizing gates * Update assembly.py * Test for optimizing gated loads * Cleanup * Fix NEG needed before if * Remove unused parameters * Update assembly.py * Fix for cachable gone --------- Co-authored-by: oz <oz@oz-MS-7B86.NAT.gliwice.vectranet.pl> Co-authored-by: chenyu <chenyu@fastmail.com>	2024-05-13 10:14:01 -07:00
George Hotz	b660f60125	all uops are now cachable (#4564 ) * all uops are now cachable * cachable is gone	2024-05-12 22:34:35 -07:00
George Hotz	02327b8adf	simple stuff from new_uops branch (#4563 )	2024-05-12 22:18:05 -07:00
George Hotz	347a3acb37	add renderer class (#4524 ) * add renderer class * tests pass * fix pylint * fix tensor cores	2024-05-10 21:40:02 -07:00
George Hotz	4eef1ee9bf	move renderer into options (#4514 ) * move renderer into options * fix tests * renders are functions	2024-05-10 10:01:51 -07:00
George Hotz	cb7289f9c9	remove clang program header (#4422 ) * remove clang program header * proper max * bools are numbers * fix compile enet	2024-05-04 08:38:01 -07:00
George Hotz	f635c4d273	fix define global (#4383 ) * fix define global * remove name from DEFINE_GLOBAL * fix fuzzing * fix ptx * fix python	2024-05-01 22:32:56 -04:00
Szymon Ożóg	f1ebcffb87	Ptx beam fix (#4296 ) * Fix beam search for PTX * fix ptr arm test	2024-04-25 15:39:39 -04:00
Szymon Ożóg	002a14088e	Ptx store gate cast to bool (#4284 ) * Cast gate to bool * Update * Add PTX fuzzing to benchmark	2024-04-24 11:43:44 -04:00
Szymon Ożóg	6c25f1abf7	Optimize ptx loops (#4263 ) * Optimize PTX loops * Update assembly.py	2024-04-23 12:20:14 +04:00
chenyu	1ef9c50fd7	Update ssa input order and annotate types in cstyle and assembly (#4117 ) variable prefix is never optional (removed the default "t") and UOp can be optional (added the default None).	2024-04-09 13:10:29 -04:00
George Hotz	e4a1858471	revert command queue (#4097 )	2024-04-06 08:58:18 -07:00
chenyu	5e6e6c9a67	use ConstType in various const function type hint (#4074 )	2024-04-04 20:32:07 -04:00
Szymon Ożóg	ba118abfec	improved caching for pointer arithmetics in ptx (#3922 ) * improved caching for pointer arithmetics * Add test for pointer arithmetics caching * Refactor test	2024-04-04 07:33:48 -07:00
Szymon Ożóg	68fe3527f1	Tensor core ptx (#3894 ) * tensor cores * Merge from master * faster program start in llvm (#3897) * Fix the result permutation in einsum (#3895) * Fix permutation of result indices in einsum. * Delete stray line used for breaking tests * Fix linter error by renaming twice-used variable --------- Co-authored-by: chenyu <chenyu@fastmail.com> * touchup einsum (#3900) don't need rhs_letters * hotfix check ckpts before writing achieved model (#3901) this killed tinybox green run * replace dtype.name str with render_dtype (#3903) fixed some bf16 cast issue since it does not have `.name`. also more robust if there are lang specific type override * add --minimal flag to nvrtc (#3899) * wmma: fix the AMD TC threads to split the first 16 threads (#3904) previously it was incorrectly aliasing 16 into the size 8 upcast on the store alias. now it splits it properly into 8 and the remaining 2 into the correct local stride * training cifar with BF16 on CUDA (#3905) * training cifar with BF16 on CUDA memory usage is between float and half due to numpy calls on dataset preprocessing, which converts into float. * simpler bf16 functions * bf16 cifar works for HSA too just very slow * simpler bf16 functions, we love cuda * include negative float in test_dtype (#3884) * include negative float in test_dtype * that is ub * too annoying * pack can overflow * add to benchmark * change var name to satisfy mypy * spacing * Update to new TensorCore format * Spacing --------- Co-authored-by: nimlgen <138685161+nimlgen@users.noreply.github.com> Co-authored-by: Alejandro F Queiruga <33233447+afqueiruga@users.noreply.github.com> Co-authored-by: chenyu <chenyu@fastmail.com> Co-authored-by: sekstini <127142660+sekstini@users.noreply.github.com> Co-authored-by: Francis Lam <flam@alum.mit.edu> Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com>	2024-04-04 07:32:31 -07:00
Szymon Ożóg	92378fb5b6	Ptx mulacc (#3937 ) * mulacc * Move more stuff to pattern matcher * disable callable from the == check * disable function passing in pattern matcher * Add set of dtypes pattern matching + refactor mulacc pattern	2024-04-04 00:15:25 -07:00
Szymon Ożóg	e5a9bff899	Add pattern matcher tests, move uop transforms from assembly to pattern (#4056 ) matcher	2024-04-03 09:06:43 -07:00
Szymon Ożóg	ccf3c16d6a	Refactor the use of pattern matcher in ptx (#4043 )	2024-04-02 14:19:51 -07:00
Szymon Ożóg	31c8ba8b84	Move transformations to PatternMatcher + clean up existing patterns (#3997 )	2024-03-29 19:42:39 -07:00
chenyu	d9ff636cf5	use is to compare with enum (#3993 ) * use is to compare with enum currently it's mixed between `==` and `is`, moved all to `is` * more	2024-03-29 13:02:56 -04:00
chenyu	80116be9a5	for loop to generate hip math functions for different floats (#3967 ) * for loop to generate hip math functions for different floats * slightly nicer	2024-03-27 23:24:29 -04:00
Francis Lam	7c5729a3bd	wmma: refactor to remove wmma_func and create TC funcs as needed (#3945 ) * wmma: refactor to remove wmma_func and create TC funcs as needed * test_linearizer: disable bf16 CUDA during emulation testing * cstyle: clean up creation of CUDA vec dtypes * extra/gemm: add option to accumulate to bfloat16 * cleanups * benchmark: add CUDA bfloat16 matmul * more cleanups	2024-03-27 16:43:09 -04:00
chenyu	88b24df40a	touchup remove `float()` in cstyle render_const for float64 (#3962 )	2024-03-27 16:08:28 -04:00
chenyu	6c7df1445b	enforce UOps.CONST arg has python type based on dtype (#3952 ) added an assert in uops, remove the cast in renderer	2024-03-27 01:41:38 -04:00
chenyu	ef537672bf	bf16 support in metal (#3929 ) it runs if device gpu supports bfloat. updated ci benchmark too	2024-03-25 23:17:36 -04:00
Arseny Kapoulkine	514c43201d	Fix issues with pointer provenance in load/store through ALU (#3916 ) * Track pointer provenance in load/store through ALU Previously load/store could be incorrectly rendered into ld.global/st.global when the input was an ALU op that performed an address computation with DEFINE_LOCAL on one of the arguments. * Simplify the load provenance workaround The issue is that we can render the same code twice, and on the second run the opstream is already modified so that vin[0] isn't a DEFINE_, which overwrites initially correct .shared wth .global. Add a couple tests for basic local use * Skip local tests on LLVM since it doesn't implement DEFINE_LOCAL	2024-03-25 14:41:05 -07:00
Arseny Kapoulkine	715850aef9	Fix sm89 PTX=1 compilation (#3915 ) * Fix sm89 PTX=1 compilation The minimum PTX version that supports sm89 is 7.8 (same version also supports sm90); without this ptxas fails when running tinygrad with PTX=1 on RTX 4090. * Use int(arch[3:]) for forward compat with SM10.0 if that happens	2024-03-24 20:32:29 -07:00
Szymon Ożóg	2d0bfdf01c	ptx cleanup (#3893 )	2024-03-24 14:54:45 -07:00
chenyu	e22d78b3d2	training cifar with BF16 on CUDA (#3905 ) * training cifar with BF16 on CUDA memory usage is between float and half due to numpy calls on dataset preprocessing, which converts into float. * simpler bf16 functions * bf16 cifar works for HSA too just very slow * simpler bf16 functions, we love cuda	2024-03-24 01:37:47 -04:00
chenyu	a2b2597fc2	replace dtype.name str with render_dtype (#3903 ) fixed some bf16 cast issue since it does not have `.name`. also more robust if there are lang specific type override	2024-03-23 19:25:48 -04:00
chenyu	18e0cef14d	cheap less lines in ptx (#3890 ) enought to merge lars	2024-03-23 01:12:31 -04:00
George Hotz	0c197b9cf3	hotfix: hip bfloat formatting	2024-03-22 11:52:05 -07:00
chenyu	c5467e5bd6	diverse test value in test_dtype DATA based on dtype (#3864 ) * diverse test value in test_dtype DATA based on dtype * eh fix typo * that too? * PTX does not support i8 and s8 * skip that * unused line * pus the hack back * remove that	2024-03-22 14:22:06 -04:00
Szymon Ożóg	624bc89910	PTX - implement float 4, ptr arithmetics and other speed improvements (#3775 ) * ptx float4 implementation * remove from cache when trimming uops * Gate for float4 * Linting fix * disable test reasonable time for ptx * import getenv * Update uops.py * linter * Add div test for half * upcast if op does not support operation * fix offset * Run only if dtype supported * zero out registers when accessing by pred + cleanup * Remove trailing whitespace * revert * spacing fix * move cache clearing outside loop * did this suddenly start working? * unused import removed * Remove cast * Use pattern matching * linting --------- Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com>	2024-03-22 08:54:02 -07:00
George Hotz	f4055439dc	don't include hip common (#3851 ) * don't install hip common * only that * Revert "only that" This reverts commit `85f22015d9`. * less * needed * sep comgr * header file * 6.0.2 * update hsa * hsakmt * Revert "hsakmt" This reverts commit `d3a118078e`.	2024-03-22 08:50:50 -07:00
qazal	4a27ce6ec9	tiny version of amd_hip_bfloat16 (#3868 ) * add src_dtype * add maker * add bfloat16 * simpler	2024-03-22 08:37:30 -07:00
nimlgen	6b8c66e04f	fix broken loops in llvm (#3751 )	2024-03-15 11:57:51 +03:00
chenyu	75d4344cda	UOps.BITCAST (#3747 ) * UOps.BITCAST implicitly fixed no const folding for bitcast * python backend * ptx * consistent llvm	2024-03-14 21:00:35 -04:00
George Hotz	f1dd8928c9	where fold prereqs (#3718 )	2024-03-13 10:01:43 -07:00
George Hotz	3af1c1051a	Revert "bring reciprocal back (#3687 )" (#3692 ) This reverts commit `bcf6fbd3b2`.	2024-03-11 15:55:14 -07:00
George Hotz	ef44c8959b	Revert "rewrite recip to div (#3690 )" (#3691 ) This reverts commit `2b089bfd18`.	2024-03-11 15:54:58 -07:00
George Hotz	2b089bfd18	rewrite recip to div (#3690 ) * rewrite recip to div * fix bug in uops add	2024-03-11 15:52:24 -07:00
George Hotz	bcf6fbd3b2	bring reciprocal back (#3687 ) * bring reciprocal back * better * explicit dtype for recip * llvm tighter * sigmoid can use RECIP	2024-03-11 14:19:54 -07:00
George Hotz	0f16729023	RDNA3: restore launch bounds (#3672 ) * bring launch bounds back * works * that second flag didn't do anything * fix linter	2024-03-10 10:27:52 -07:00
chenyu	d7452c2a20	clean up llvmir builder (#3671 ) ``` _block -> block builder._block.module -> builder.module var_dtype -> dtype ```	2024-03-09 21:19:36 -05:00
George Hotz	6e50582e62	working to improve ptx (#3647 ) * working to improve ptx * fix compile fail	2024-03-07 12:39:31 -08:00

1 2 3 4 5

204 Commits