tinygrad

mirror of https://github.com/tinygrad/tinygrad.git synced 2026-01-09 15:08:02 -05:00

Author	SHA1	Message	Date
George Hotz	efc99d0c55	assembly/amd: more refactors (#13907 ) * assembly/amd: more refactors * more refactors * more refactors * simpler emu * generate.py * regen all * cleanups * more * work * more readme * lil	2025-12-30 16:13:24 -05:00
George Hotz	81cf9ea0ab	rename to extra.assembly.amd (#13879 )	2025-12-29 14:10:55 -05:00
George Hotz	25ef866e89	write python emulator from RDNA3 psuedocode in pdf (#13841 ) * write python emulator from RDNA3 psuedocode in pdf * emu2 * more emu * working * more psueod * progress * cleanups * delete junk * delete stale files * just emu * work * emu compare * bemu * cleanups and more failures * revert bench emu * fix emu cmp * four tests fail * bugfixes * dsl * ext * refactor * dsl * div scale fix * test_emu * fix emu tests * pcode * test pcode * top imports * fix test_emu to use run_asm * emu tests on real hardware * more tests * more emu tests * more * work * work * bug fix * bugfixes * fix fp16 gemm * all ops tests pass in emulator * fix llvm tests * fix a few more tests * fix mockgpu timeout	2025-12-29 07:39:53 -05:00
nimlgen	c6769badc2	mockgpu: async support (#13868 ) * mockgpu: async support * cpu	2025-12-29 13:18:37 +03:00
George Hotz	9d94b8c6b2	python asm dsl in extra + python REMU (#13436 ) * having fun with python asm dsl * rdna3 * meh * all in rdna3 * work * more work * work * integration * tests * simpler * simpler * asm * better * simpler * progress * emu * simpler * emu * tests * types * vopd * cleaups * work * memory ranges * add tracing * refactors * run_asm exit * more readable * compare to remu * test gemm * bug + stale * more tests * refactor * tests fix * more ins * more instructions * refactor * faster * match case * match case * simpler * work * tests * run_asm * work * bug fixes * more emu * alu/emu * refactor * no pipeline emu yet * alu direct * fix * bugfixes + new test * fix exceptions in emulators * update gen.py * pylint * no pdf * improve bench_emu * speedups * cleanups * more tests	2025-12-25 13:04:14 -05:00
nimlgen	455dd88236	nv: minimal hevc (#13502 ) * nv: minimal hevc * validate * not needed * tralin * var * cpu * fxi * desc * move * cleanup	2025-11-30 16:46:55 +03:00
nimlgen	9bb17c53ea	amd: timer fix (#13267 )	2025-11-17 13:59:03 +03:00
nimlgen	3e63831b98	nv: support 580+ drivers (#13269 ) * nv: 580+ support * start * f * fake * fix	2025-11-14 21:44:16 +08:00
nimlgen	14eb48b13a	autogen: rename nv_gpu to nv_570 (#13273 ) * autogen: rename nv_gpu to nv_570 * rename	2025-11-14 20:07:19 +08:00
nimlgen	f72b1fbca4	nv: read numClasses (#13271 ) * nv: read numClasses * fix * d	2025-11-14 19:43:25 +08:00
Christopher Milan	09f3aae169	In-tree autogen: all C libraries (#13220 ) * checkout files from autogen branch * ioctl with payload * fix am generations * properly fix generations This reverts commit `b2a54f4f41`. * revert discovery.h * support pragma pack(1) * typo * better getter * typo * NVCEC0_QMDV05_00_RELEASE[01]_ENABLE * align support * anon handling fix --------- Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com>	2025-11-13 18:57:44 -08:00
nimlgen	4b001ec723	amd: pmc in mockgpu (#13000 ) * amd: pmc in mockgpu * fix * do not open in ci	2025-10-30 01:52:02 +08:00
chenyu	585bd95b50	fix ruff 0.14.0 [pr] (#12547 )	2025-10-09 01:52:30 -04:00
nimlgen	4a756a37d8	amd: support rocm7 (#12502 ) * amd: support rocm7 * mock	2025-10-08 14:30:39 +08:00
nimlgen	c6e342cdac	mockgpu: no hang if gpuocelot failed (#11915 )	2025-08-30 00:44:49 +03:00
nimlgen	fa695ac1ce	ci: mac gpuocelot (#11906 ) * gm * fix? * ops * imp * xx * add file	2025-08-28 23:29:43 +03:00
nimlgen	9c9e337c78	amd: parse soc enums (#11727 ) * amd: parse soc enums * remove from mock * fix * minimal amd_gpu	2025-08-19 15:06:09 +03:00
nimlgen	4176b24264	amd: support xcc in regs (#11670 ) * amd: support xcc in regs * mockamd * typong	2025-08-14 21:20:11 +03:00
uuuvn	c29c46853f	Very basic mock sqtt (#10512 ) This mockgpu sqtt emulation will just ignore basically everything and end up with a 0x1000 size trace full of zeroes, but just testing for things like register rename is better than nothing i guess	2025-05-26 14:38:28 -07:00
uuuvn	27c12be471	amd mockgpu graph support (#10385 ) For testing remote graph stuff (prompted by #10371) in ci	2025-05-18 09:43:16 -07:00
wozeparrot	f59ecf2116	fix: mockgpu cuda timing (#10343 )	2025-05-15 14:14:14 -07:00
nimlgen	0fbe494c6b	usb: cache writes into 0xa000 (#10191 ) * usb: cache writes into 0xa000 * mock * match parent spec * ugh	2025-05-07 16:03:35 +03:00
nimlgen	685d5c46df	usbgpu: send pci write in batches (#10190 ) * usbgpu: send pci write in batches * mock	2025-05-07 14:41:56 +03:00
nimlgen	30bd6a619f	usb gpu (#8766 ) * start gpu * progress * fixes * read correct * libusb * libusb works * support asm24 * hmm * one access file * fix extra * start AMBar * works on am * back to usb * patch fw * full fast write into a bar * ugh, minus one gpus, next please * mute libusb for now * usb for asm24 * 63 * hmm * ops * rescan * and gpu shoudl be there * enumerate them? * usbgpu bus 4, 100% reliable (draft) * lil * works * comments * add DEBUG * cleaner * simplest * Revert "simplest" This reverts commit `1d00354c16`. * Revert "cleaner" This reverts commit `c5662de956`. * assert we find gpu * that's simpler * this back * simpler? * correcT * work * nonsense * works with more checks * this works * the 6s in the right place * reliable now * fix after reboot * set config * 1s timeouts * close to fw loading * streams * usbhub works * endpoints * fix * want to test tiny10 * move to tiny 10 * fix gpu * ugly speed * smth * mostly broken, but signals and dmas * do not reset gpu every time * changes to run kernels * ugh, not working * t10 * pg and sc files * some prog * um? * somehow it works * patched for 24 * some tries * minimal * moving * back to working * so sloooooow * move to controller * usb.py rewrite * rework * cleaner 1 * cleaner 2 * cleaner 3 * new abstractions * aft merge * init controller * cleaner 4 * cleaner 5 * patcher + tiny changes * ignore that * cleaner 6 * after rebase * cleaner 7 * bring it back * start linter war * linter 2 * autogen was missing * fix autogen * typing * better? * mypy * extra/legacy rename and cleaner * shuffle * better printing * tiny changes and tests --------- Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com>	2025-05-01 18:03:47 +03:00
nimlgen	db51133537	rename HWInterface -> FileIOInterface (#9989 ) * rename HWInterface -> FileIOInterface * ugh	2025-04-22 22:18:57 +03:00
nimlgen	bd580d8ea4	hcq: use mmio interface in nv (#9986 ) * hcq: start mmio interface * allow double cast * revert * faster? * simpler, not needed more now * dd * types * fix	2025-04-22 21:58:12 +03:00
qazal	16dfe0a902	upstream remu (#9921 )	2025-04-18 01:57:36 +03:00
uuuvn	dd9aae02c3	Refactor ops_amd.py (MI300X prereq) (#9428 )	2025-03-29 00:17:20 +07:00
b1tg	2c87a22cf2	fix prg size calculation when there are adjacent mapped ranges (#9498 ) Co-authored-by: b1tg <b1tg@users.noreply.github.com>	2025-03-19 11:55:03 +08:00
uuuvn	b75f307234	amd: autogen ip bases (#9360 )	2025-03-05 22:30:38 +03:00
nimlgen	70db8c3003	hcq: dyn alloc signals (#9238 ) * hcq: dyn alloc signals * types and uniqueue devs * typing * mypy * mypy one more time * test * make fds to not intersect in mockgpu between drivers	2025-02-25 17:22:24 +03:00
JaSpa99	d2ff55e9c6	OSX GPUOcelot (#8209 ) * add patches * add osx test in ci * macos specific uvm, gpfifo mask * only do that for now * Revert "add patches" This reverts commit `80d3112a57`. * use fork for now * workflow only one worker * merge osxtests with tests * Revert "merge osxtests with tests" This reverts commit `3461c8f46c`. * macos pagesize 16384 --------- Co-authored-by: nimlgen <138685161+nimlgen@users.noreply.github.com> Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com>	2025-02-13 12:24:29 +08:00
nimlgen	166670a2f2	nv: fill grid/block sizes (#9025 )	2025-02-11 16:30:30 +03:00
nimlgen	88add71c25	amd: increase sdma copy size (#8989 ) * amd: increase sdma max copy size * rm this * fix * fx * ops	2025-02-09 20:53:35 +03:00
nimlgen	e5a3f60fc2	am: remove libpciaccess dep (#8980 ) * am: remove libpciaccess dep * offset in mockhwiface * op * fake regions	2025-02-09 16:06:55 +03:00
nimlgen	d224d0ed7f	nv: fix fault info (#8587 ) * nv: fix fault info * and emu for amd * skip if not mock	2025-01-13 14:38:43 +03:00
nimlgen	38b5ac4d4a	mypy for mockgpu/cuda & dsp/run (#8575 )	2025-01-12 18:25:39 +03:00
nimlgen	aa3d612df2	add script to install amd mockgpu on macOS (#8536 ) * upload artifact every time * hm * sh script * hm * hm2 * hm2 * hm2 * no sudo * def paths * small comments * text * try auth for bigger limits	2025-01-09 01:29:25 +03:00
patrini32	afef69a37d	MOCKGPU on mac os (#8520 ) * tweaks for macos * fix * fix * typo * remove nvidia changes * remove nv related changes * change address back	2025-01-07 20:27:43 +03:00
nimlgen	ab3ac2b58d	hw interface abstraction (#8524 ) * use HWInterface in autogen * mockgpu * HWInterface * more HWInterface * fix * fix * old code * fix * implicit field definition * add offset check to mockgpu too * refactor * forgot to pass flags + read rewrite * test * play with vfio * nv: this should be kept * try this * vfio * rm overwrite=True * linetr * do not reinit kfd * minor * mypy * mock * init them once --------- Co-authored-by: patrini32 <patrini23@proton.me>	2025-01-07 18:18:28 +03:00
nimlgen	9bc317d5d2	mockcuda (#8503 ) * init mockcuda * run gpu ocelot * fix * sfixes * disable broken tests * linter * these fails as well * pylint * myypy * this fails on real platforms as well * mypy please	2025-01-05 01:23:57 +03:00
nimlgen	c18307e749	AM driver (#6923 ) * connect to gpu * rlc init? * gfx comp start init * early init is hardoded, some progress with fw * gart * progress, next mqd * ring setup, still does not execute anything * ugh write correct reg * pci2: vm * pci2: start psp * vm seems to work * pci2: gfx start * pci2: fix psp ring resp * pci2: try ring * pci2: mes and some fixes * pci2: some progress * pci2: progress * pci2: mm * pci2: discovery * pci2: correct apertures * pci2: b * pci2: i * pci2: l * pci2: o * pci2: cmu * pci2: mes_kiq works * pci2: mes * pci2: kcq does not work( * pci2: unhalt gfx * ops_am * minor * check if amdgpu is there, or we will crash * bring back graph, it just works * less prints * do not init mes (not used) * remove unused files * ops_am: start move into core * ops_am: works * clcks, but still slower * faster + no mes_kiq * vm frags + remove mes * cleanup fw * gmc tiny cleanup * move to ops_amd * comment out what we dont really need * driverless * close in speed * am clean most of ips * gmc to ips * cleaner * new vm walker * comment old one * remove unsued autogens * last write ups * remove psp hardcoded values * more * add logs * ih * p2p and sdma * vfio hal and interrupts * smth * amd dev iface * minor after rebase * bind for sdma * Revert "bind for sdma" This reverts commit `a90766514d`. * tmp * debug new mm * ugh, allreduce hangs fixed * p1 * works * no pci.py * cleaner a bit * smth * tiny cleanups * cleaner a bit * pciiface * linter * linter 2 * linter 3 * linter * pylint * reverted unrelated changes * unrelated * cmp tool * ugh wrong fw * clockgating * unrelated * alloc smaller chunks * this * opt sigs * collect stat * ops * upd * proclogs * proclogs2 * vfio * ruff * linter pylint * oops * mypy p1 * mem fix * mypy p2 * mypy p3 * mypy p4 * correct * minor * more tests * linter in tests * pci_regs header * minor write up * setup * do not require libs --------- Co-authored-by: George Hotz <72895+geohot@users.noreply.github.com>	2024-12-31 23:06:17 +03:00
qazal	6422936b62	fix pre-commit ruff error [pr] (#8405 )	2024-12-26 00:12:57 +08:00
nimlgen	a647f3dd2c	move mockgpu to tests [pr] (#8396 ) * move mockgpu to tests * linter * i'm so sorry * sorry, python * path	2024-12-24 23:48:02 +03:00

44 Commits