[MTIA] Allow users who know what they are doing to ignore all device mismatches in tracing and take a preferred device. by patrick-toulme · Pull Request #159931 · pytorch/pytorch

patrick-toulme · 2025-08-06T04:14:27Z

Summary:
Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical.

Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations.

If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors.

Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing.

Users can set

  torch._functorch.config.fake_tensor_prefer_device_type = 'mtia'

to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA.

Test Plan:
Added two unit tests.

Rollback Plan:

Differential Revision: D79698438

pytorch-bot · 2025-08-06T04:14:31Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/159931

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki or our office hours

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit db686a4 with merge base 3a2c3c8 ():

UNSTABLE - The following job is marked as unstable, possibly due to flakiness on trunk:

pull / linux-jammy-py3_9-clang9-xla / test (xla, 1, 1, lf.linux.12xlarge, unstable) (gh) (#158876)
/var/lib/jenkins/workspace/xla/torch_xla/csrc/runtime/BUILD:476:14: Compiling torch_xla/csrc/runtime/xla_util_test.cpp failed: (Exit 1): gcc failed: error executing CppCompile command (from target //torch_xla/csrc/runtime:xla_util_test) /usr/bin/gcc -U_FORTIFY_SOURCE -fstack-protector -Wall -Wunused-but-set-parameter -Wno-free-nonheap-object -fno-omit-frame-pointer -g0 -O2 '-D_FORTIFY_SOURCE=1' -DNDEBUG -ffunction-sections ... (remaining 229 arguments skipped)

This comment was automatically generated by Dr. CI and updates every 15 minutes.

facebook-github-bot · 2025-08-06T04:14:38Z

This pull request was exported from Phabricator. Differential Revision: D79698438

patrick-toulme · 2025-08-06T04:21:42Z

I have seen this FakeTensorDevicePropagation in many issues. We should allow users an escape hatch to get around any intermediate device mismatch if they are confident the mismatch is in intermediate (non materialized) tensors and not graph inputs.

#144748
#151670
#151296

I have also seen this issue many times internally.

…mismatches in tracing and take the non CPU device. (pytorch#159931) Summary: Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_non_cpu_device = True ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438

facebook-github-bot · 2025-08-06T04:26:40Z

This pull request was exported from Phabricator. Differential Revision: D79698438

…mismatches in tracing and take the non CPU device. (pytorch#159931) Summary: Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_non_cpu_device = True ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438

facebook-github-bot · 2025-08-06T04:34:13Z

This pull request was exported from Phabricator. Differential Revision: D79698438

patrick-toulme · 2025-08-06T15:35:02Z

Tests passed local but failed on PR. Will debug tests and push again

…mismatches in tracing and take the non CPU device. (pytorch#159931) Summary: Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_non_cpu_device = True ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438

facebook-github-bot · 2025-08-06T15:59:47Z

This pull request was exported from Phabricator. Differential Revision: D79698438

…mismatches in tracing and take the non CPU device. (pytorch#159931) Summary: Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_non_cpu_device = True ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438

facebook-github-bot · 2025-08-06T16:06:28Z

This pull request was exported from Phabricator. Differential Revision: D79698438

torch/_functorch/config.py

…mismatches in tracing and take the non CPU device. (pytorch#159931) Summary: Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_device_type = 'mtia' ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438

facebook-github-bot · 2025-08-06T22:51:06Z

This pull request was exported from Phabricator. Differential Revision: D79698438

…mismatches in tracing and take the non CPU device. (pytorch#159931) Summary: Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_device_type = 'mtia' ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438

facebook-github-bot · 2025-08-06T22:52:43Z

This pull request was exported from Phabricator. Differential Revision: D79698438

…mismatches in tracing and take the non CPU device. (pytorch#159931) Summary: Pull Request resolved: pytorch#159931 Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_device_type = 'mtia' ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438

facebook-github-bot · 2025-08-06T22:55:58Z

This pull request was exported from Phabricator. Differential Revision: D79698438

patrick-toulme · 2025-08-07T22:31:23Z

@pytorchbot merge

pytorchmergebot · 2025-08-07T22:33:13Z

Merge failed

Reason: This PR needs a release notes: label
If your changes are user facing and intended to be a part of release notes, please use a label starting with release notes:.

If not, please add the topic: not user facing label.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "topic: not user facing"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

Details for Dev Infra team

Raised by workflow job

patrick-toulme · 2025-08-07T22:34:41Z

@pytorchbot label "topic: not user facing"

patrick-toulme · 2025-08-07T22:35:02Z

@pytorchbot merge

pytorchmergebot · 2025-08-07T22:36:54Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

…mismatches in tracing and take a preferred device. (pytorch#159931) Summary: Device mismatches in tracing can most often be ignored. These are only logical mismatches not physical. Take any intermediate computation, and that computation will not actually materialize in a compiled binary execution. So a device mismatch in the middle of the program is not real. The runtime will never materialize those tensors on CPU device during the execution, as they are temporary allocations. If a user knows his tensors at graph input are all on the correct device, then he can ignore all tracing errors. Users who know what they are doing should have an escape hatch to ignore any device mismatch in tracing. Users can set ``` torch._functorch.config.fake_tensor_prefer_device_type = 'mtia' ``` to forcefully override any mismatch and prefer the non cpu device. This unblocks vLLM graph mode for MTIA. Test Plan: Added two unit tests. Rollback Plan: Differential Revision: D79698438 Pull Request resolved: pytorch#159931 Approved by: https://github.com/jansel

pytorch-bot bot added the ciflow/inductor label Aug 6, 2025

facebook-github-bot added the fb-exported label Aug 6, 2025

patrick-toulme force-pushed the export-D79698438 branch from 342a96b to da11020 Compare August 6, 2025 04:26

patrick-toulme force-pushed the export-D79698438 branch from da11020 to 5759a38 Compare August 6, 2025 04:34

patrick-toulme force-pushed the export-D79698438 branch from 5759a38 to ad78a94 Compare August 6, 2025 15:59

patrick-toulme force-pushed the export-D79698438 branch from ad78a94 to 874c000 Compare August 6, 2025 16:06

patrick-toulme requested review from aorenste, jansel, laithsakka and masnesral and removed request for aorenste and masnesral August 6, 2025 16:16

jansel requested changes Aug 6, 2025

View reviewed changes

torch/_functorch/config.py Outdated Show resolved Hide resolved

patrick-toulme force-pushed the export-D79698438 branch from 874c000 to 650c282 Compare August 6, 2025 22:43

patrick-toulme force-pushed the export-D79698438 branch from 650c282 to 7209374 Compare August 6, 2025 22:43

patrick-toulme force-pushed the export-D79698438 branch from 7209374 to 79c3d75 Compare August 6, 2025 22:50

patrick-toulme force-pushed the export-D79698438 branch from 79c3d75 to 69f70de Compare August 6, 2025 22:51

patrick-toulme force-pushed the export-D79698438 branch from 69f70de to 5961c3b Compare August 6, 2025 22:52

patrick-toulme requested a review from jansel August 6, 2025 22:54

patrick-toulme force-pushed the export-D79698438 branch from 5961c3b to db686a4 Compare August 6, 2025 22:56

jansel approved these changes Aug 7, 2025

View reviewed changes

patrick-toulme removed request for laithsakka and masnesral August 7, 2025 22:29

pytorchmergebot added the merging label Aug 7, 2025

pytorchmergebot removed the merging label Aug 7, 2025

pytorch-bot bot added the topic: not user facing topic category label Aug 7, 2025

pytorchmergebot added the merging label Aug 7, 2025

pytorchmergebot closed this in d46768d Aug 7, 2025

pytorchmergebot added Merged and removed merging labels Aug 7, 2025

Conversation

patrick-toulme commented Aug 6, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

pytorch-bot bot commented Aug 6, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/159931

✅ You can merge normally! (1 Unrelated Failure)

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

patrick-toulme commented Aug 6, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

patrick-toulme commented Aug 6, 2025

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

facebook-github-bot commented Aug 6, 2025

Uh oh!

patrick-toulme commented Aug 7, 2025

Uh oh!

pytorchmergebot commented Aug 7, 2025

Merge failed

Uh oh!

patrick-toulme commented Aug 7, 2025

Uh oh!

patrick-toulme commented Aug 7, 2025

Uh oh!

pytorchmergebot commented Aug 7, 2025

Merge started

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

patrick-toulme commented Aug 6, 2025 •

edited

Loading

pytorch-bot bot commented Aug 6, 2025 •

edited

Loading

patrick-toulme commented Aug 6, 2025 •

edited

Loading