tinygrad

mirror of https://github.com/tinygrad/tinygrad.git synced 2026-01-25 06:48:22 -05:00

Author	SHA1	Message	Date
nimlgen	b6937acb7e	fix casting behavior for interpreted buffers (#1525 )	2023-08-13 19:21:37 -07:00
David Heidelberg	13659ac6fa	examples: numpy() array returns only one value, not an array (#1534 ) Fixes issue: ``` loss_cpu = loss.detach().numpy()[0] ~~~~~~~~~~~~~~~~~~~~~^^^ IndexError: too many indices for array: array is 0-dimensional, but 1 were indexed ``` Signed-off-by: David Heidelberg <david@ixit.cz>	2023-08-13 14:33:05 -07:00
chenyu	3e0c2d256f	symbolic shapetracker (#1506 ) * symbolic shapetracker * no need * keep only symbolic and clean up * explicit // and % Node support * NumNode * Node	2023-08-12 12:22:58 -07:00
Pavol Rusnak	875da762a8	fix file race condition in ops_clang (#1458 )	2023-08-12 09:31:46 -07:00
JaSpa99	d3d58a37e5	Bert: use Tensor.scaled_dot_product_attention (#1528 ) * use scaled attn from Tensor * add a test for bert * linter * no more tokenizer * without loading weights * remove prints * tribute to linter lords * smaller input and less runs * small bert	2023-08-12 08:46:04 -07:00
Szymon Ożóg	330fb7b1a3	Print more meaningfull hip error messages (#1530 )	2023-08-12 07:16:20 -07:00
wozeparrot	29d5801387	distributed collectives (#1519 ) * feat: world * feat: tests * feat: no more backwards * feat: recv into * feat: whoops * feat: test in ci * feat: some debug logging * feat: workflow naming * feat: need to set pythonpath * feat: just send to same device * feat: allreduce * feat: test * feat: need contiguous * feat: test in ci * feat: exit with correct code * feat: don't need that * feat: opencl wait_for just doesn't work * feat: synchronize on out * feat: try? * feat: try again? * feat: add extra realizes * feat: print * feat: seed * feat: tol * feat: test ones and zeros * feat: remove print * feat: are you just flaky * feat: seperate scatter and gather? * feat: just try synchronizing * feat: remove print again * feat: bring back difference * feat: no sync * feat: revert that * feat: back to wait_for * fix: typo	2023-08-11 10:22:07 -07:00
Jacky Lee	2e85fce068	Transformer: use Tensor.scaled_dot_product_attention (#1520 )	2023-08-11 09:00:37 -07:00
George Hotz	38fe84d92b	cleanup mlops (#1521 ) * cleanup mlops * that line belongs there	2023-08-10 19:53:28 -07:00
George Hotz	47f18f4d60	[New] SD: Refactor AttnBlock, CrossAttention, CLIPAttention to share code (#1516 ) (#1518 ) * Refactor AttnBlock, CrossAttention, CLIPAttention to share code * Reshape and transpose in loop * Bugfix on attention mask Co-authored-by: Jacky Lee <39754370+jla524@users.noreply.github.com>	2023-08-10 15:04:18 -07:00
wozeparrot	7e7c9001e9	distributed world (#1481 ) * feat: world * feat: tests * feat: no more backwards * feat: recv into * feat: whoops * feat: test in ci * feat: some debug logging * feat: workflow naming * feat: need to set pythonpath * feat: just send to same device	2023-08-10 10:00:51 -07:00
George Hotz	e3c6c0c6db	add GPT2 example (#1511 ) (#1514 ) * add gpt2 to examples * some cleanup * fixes * argparse + scaled_dot_product_attention * add timing * add to benchmark Co-authored-by: YassineYousfi <yassine.y10@gmail.com>	2023-08-10 09:09:47 -07:00
George Hotz	c82bd59b85	Revert "SD: Refactor AttnBlock, CrossAttention, CLIPAttention to share code (#1513 )" (#1515 ) This reverts commit `85e02311a2`.	2023-08-10 09:08:51 -07:00
Jacky Lee	85e02311a2	SD: Refactor AttnBlock, CrossAttention, CLIPAttention to share code (#1513 ) * Refactor AttnBlock, CrossAttention, CLIPAttention to share code * Reshape and transpose in loop	2023-08-10 08:52:33 -07:00
geohotstan	07b79f210f	llvmir support for bool <-> float casting (#1492 )	2023-08-09 13:12:52 -04:00
wozeparrot	351684395c	dont run on fork (#1510 )	2023-08-09 13:06:45 -04:00
wozeparrot	88e2e0c8a3	Revert "don't try to run benchmark on forks" (#1508 )	2023-08-09 12:59:49 -04:00
wozeparrot	65b65b760b	don't try to run benchmark on forks (#1507 )	2023-08-09 12:59:19 -04:00
George Hotz	c417cd3c97	fast HIP gemm -> 100 TFLOPS (#1476 ) * fast HIP gemm * wmma * correct b * fix spilling * 60 TFLOPS * 64 TFLOPS * 65 TFLOPS	2023-08-09 06:54:15 -07:00
David Hou	1766f0c0cf	use ConstOp for valid.max == 0 (#1501 ) * use ConstOp for valid.max == 0 * don't render valid for invalid load cache key * Update linearizer.py --------- Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com>	2023-08-09 00:01:59 -07:00
Jacky Lee	ef5f648e2f	Tensor.scaled_dot_product_attention to match torch, used in LLaMA, and tested (#1502 ) * Implement scaled_dot_product_attention and test * Support attn_mask * Support is_causal too * Use in llama * Don't forget to reshape * Set requires_grad=False for causal * Remove staticmethod * Remove extra spaces	2023-08-08 23:27:13 -07:00
nimlgen	dabfd7569a	use allclose instead of equals in test_jit (#1504 ) Closes #1503	2023-08-08 22:22:17 -07:00
chenyu	827d13e64e	correct patch JIT llama chat (#1500 )	2023-08-08 19:52:09 -04:00
Yixiang Gao	7c2ea85bb0	Raise memory limit for CIFAR test (#1499 )	2023-08-08 19:40:56 -04:00
Thiago Franco de Moraes	293a10204b	Add tinygrad.renderer to packages in setup.py (#1497 )	2023-08-08 15:51:49 -07:00
chenyu	0415a48cfc	patch JIT llama chat mode (#1496 )	2023-08-08 15:15:56 -07:00
Yixiang Gao	6480a1a180	CIFAR 94.03% (#1340 ) * add disk_tensor * fix jit * new baseline before whitening * whitening through torch * whiting done currently at 91.65% * 91.99% * clean up mixup and 92.3% * clean up 92.30% * 92.49% before searching for new hyper-parameters * fix CI * fix white space * add whitening init in test * refactor, update hyperpara, 92.72% * converting whiting to tinygrad operation * update CI kernels count for CIFAR * add pad reflect * add random crop 92.53% * update hyperpara 93% * 93.15% on docker container, need to refactor the assignment for hyper param * print out weights and bias to be separated * bias/non-bias params separated * fix whitespace * clean up * refactor hyper-param with dict * refactor lr schedular params * fix whitespace * fix cross entropy loss * fix whitespace * move opt hyp to hyp dict * minor fixup * adjust model, loss scaling * 92.74% while using half of compute as before * update hyp for cutmix * random shuffle during batches * clean up * updating the model * update ConvGroup * disable gradients for batchnorm layer weights * whitespace * 93.92% * clean up * finally 94%git add .! * rewrite whitening to remove dependency on torch * whitespace * remove dependency on torch, 93.91% * back to 94.03% * clean up * update test_real_world	2023-08-08 15:13:24 -07:00
Roelof van Dijk	aa83a9e910	ci: fix gpuocelot build cache (#1474 ) Co-authored-by: Roelof van Dijk <roelof.van.dijk@vitestro.com>	2023-08-08 14:00:04 -07:00
George Hotz	d24f936501	just cmplt (#1493 ) * just cmplt * fix maximum * don't save, there's no backward * ugh, no slot either * eq is a scam	2023-08-08 13:58:10 -07:00
Roelof van Dijk	e2cf0f322e	[READY] ci: missing n=auto (#1486 ) * ci: missing n=auto * fix: add to commented test --------- Co-authored-by: Roelof van Dijk <roelof.van.dijk@vitestro.com>	2023-08-08 07:37:24 -07:00
Roelof van Dijk	0ce7511110	fix: is not use with a literal (#1487 ) Co-authored-by: Roelof van Dijk <roelof.van.dijk@vitestro.com>	2023-08-08 07:35:30 -07:00
nimlgen	932dad1a2b	fix cast bool->float in llvmir (#1480 ) Closes #1479	2023-08-07 21:30:51 -07:00
nimlgen	046fd7437a	use fake buffer for external_test_speed_llama.py (#1478 )	2023-08-07 22:05:44 -04:00
George Hotz	5fdd248617	don't download cifar (#1472 )	2023-08-06 21:38:59 -07:00
George Hotz	d78fb8f4ed	add stable diffusion and llama (#1471 ) * add stable diffusion and llama * pretty in CI * was CI not true * that * CI=true, wtf * pythonpath * debug=1 * oops, wrong place * uops test broken for wgpu * wgpu tests flaky	2023-08-06 21:31:51 -07:00
terafo	24933ab551	Actually flip local_max in CUDA (#1462 ) * Actually do the flip * Fixed typo --------- Co-authored-by: terafo <terafo@protonmail.com>	2023-08-06 10:35:25 -07:00
Diogo	d7d1011f1e	Add WEBGPU tests to CI (#1463 ) * webgpu tests * assert device is webgpu * missed env set * exclude failing ci tests * ignore test file * changed acc for adam test	2023-08-06 10:32:01 -07:00
George Hotz	486a9dbfd9	speed v torch (#1464 ) * speed v torch * always print * change print * torch speed tee * all exposed	2023-08-06 09:32:33 -07:00
George Hotz	2ab282bfec	run on update_benchmark too (#1460 ) * run on update_benchmark too * amd inference test * name it better * add 10 CIFAR training steps	2023-08-06 08:58:37 -07:00
terafo	3d41674b42	Fixed regression (#1447 ) Co-authored-by: terafo <terafo@protonmail.com>	2023-08-06 07:55:58 -07:00
George Hotz	d67e248d9b	simple bitcast 2 (#1445 ) * simple bitcast 2 * bc 2 * empty * Revert "empty" This reverts commit `d8ee083655`.	2023-08-06 00:30:50 -07:00
George Hotz	943b227cb1	only on push to master	2023-08-06 00:10:07 -07:00
George Hotz	2274e3e757	Fix benchmark (#1454 ) * do benchmarking * system * artifact * go * name artifact * only on push	2023-08-05 23:44:36 -07:00
George Hotz	bf21aec81f	do benchmarking (#1451 ) * do benchmarking * system * artifact * go * name artifact	2023-08-05 23:35:01 -07:00
nimlgen	1ba8ae62a1	Match Torch speed for sum reduction (#1387 ) Co-authored-by: Alexander Edwards <alex@alexedw.com>	2023-08-05 22:27:33 -07:00
chenyu	09ede08b23	simplify Node.sum aggregating (#1449 )	2023-08-05 22:19:36 -07:00
George Hotz	7fa730b506	external model benchmark test	2023-08-05 22:10:48 -07:00
chenyu	cb5dcc7b57	remove view_from_shape (#1448 )	2023-08-05 20:39:13 -07:00
Diogo	e2af95c2f8	moved global_max and local_max to LinearizerOptions also added assert for max bufs (#1446 )	2023-08-05 18:23:18 -07:00
George Hotz	7b8d06c9f1	test uops (#1444 ) * test uops * tests should pass * improve uops * precision	2023-08-05 12:35:56 -07:00

... 162 163 164 165 166 ...

10417 Commits