Add PyTorch inference benchmark Docker guide (+ CLIP and Chai-1) (#4654)

* update vLLM links in deploy-your-model.rst * add pytorch inference benchmark doc * update toc and vLLM title * remove previous versions * update * wording * fix link and "applies to" * add pytorch to wordlist * add tunableop note to clip * make tunableop note appear to all models * Update docs/how-to/rocm-for-ai/inference/pytorch-inference-benchmark.rst Co-authored-by: Leo Paoletti <164940351+lpaoletti@users.noreply.github.com> * Update docs/how-to/rocm-for-ai/inference/pytorch-inference-benchmark.rst Co-authored-by: Leo Paoletti <164940351+lpaoletti@users.noreply.github.com> * Update docs/how-to/rocm-for-ai/inference/pytorch-inference-benchmark.rst Co-authored-by: Leo Paoletti <164940351+lpaoletti@users.noreply.github.com> * Update docs/how-to/rocm-for-ai/inference/pytorch-inference-benchmark.rst Co-authored-by: Leo Paoletti <164940351+lpaoletti@users.noreply.github.com> * fix incorrect links * wording * fix wrong docker pull tag --------- Co-authored-by: Leo Paoletti <164940351+lpaoletti@users.noreply.github.com> (cherry picked from commit c3faa9670b)
2026-04-05 03:01:17 -04:00 · 2025-04-23 17:35:52 -04:00
parent d2ccd706a5
commit 311b4cd62b
7 changed files with 200 additions and 9 deletions
--- a/docs/data/how-to/rocm-for-ai/inference/pytorch-inference-benchmark-models.yaml
+++ b/docs/data/how-to/rocm-for-ai/inference/pytorch-inference-benchmark-models.yaml
@@ -0,0 +1,25 @@
+pytorch_inference_benchmark:
+  unified_docker:
+    latest: &rocm-pytorch-docker-latest
+      pull_tag: rocm/pytorch:latest
+      docker_hub_url:
+      rocm_version:
+      pytorch_version:
+      hipblaslt_version:
+  model_groups:
+    - group: CLIP
+      tag: clip
+      models:
+      - model: CLIP
+        mad_tag: pyt_clip_inference
+        model_repo: laion/CLIP-ViT-B-32-laion2B-s34B-b79K
+        url: https://huggingface.co/laion/CLIP-ViT-B-32-laion2B-s34B-b79K
+        precision: float16
+    - group: Chai-1
+      tag: chai
+      models:
+      - model: Chai-1
+        mad_tag: pyt_chai1_inference
+        model_repo: meta-llama/Llama-3.1-8B-Instruct
+        url: https://huggingface.co/chaidiscovery/chai-1
+        precision: float16