torchdim Python port by ezyang · Pull Request #160236 · pytorch/pytorch

ezyang · 2025-08-09T00:55:49Z

Stack from ghstack (oldest at bottom):

The big semantic change (and the reason for this port) is that we no longer monkeypatch Tensor with torchdim's special methods. The new algorithm for handling dispatch is that we first land in __torch_function__ and we see if a special FCD implementation needs to be dispatch to first, and if there is nothing we fallback to the standard level strategy.

Because there is no longer C binding equivalent of classes, we've condensed _C.Dim and Dim together, and similar for Tensor. This resulted in some bugs as the Python API is sometimes different from the C API. I've attempted to disambiguate these but there may still be mistakes (many early bugs were due to this problem). Dim and DimEntry are especially painful as Dim must abide by Tensor equality semantics, but is pointer equality in C (DimEntry doesn't have this problem). Another difference between C/Python that is subtle is we no longer get implicit conversions from Dim to DimEntry, this also caused some bugs.

Much of the mechanical porting work was done by claude code. I have a separate PR that deletes functorch._C, but it was useful having dim.cpp to point claude at it so I haven't done it in this PR. From a reviewing perspective, I need to re-review that I didn't forget to port anything, some noticeably missing "small" things are patched_dim_method. I am still in progress of carefully doing a side-by-side review of ports; "simplifications" from claude code were also a major source of bugs.

There are two major feature gaps in the implementation:

DelayedTensor and dot handling are not implemented yet. This should be reasonably easy, just need to do it. However, for the purposes of sharded propagation it is actually better not to reconstruct matmuls.
Splitting dimensions with an index like [x, y] doesn't work. The problem is that __getitem__ interprets this as advanced indexing and sends the list to torch.tensor to turn into a tensor, instead of being eligible for __torch_function__. I think I might need to hard code a special case for this or something?

Signed-off-by: Edward Yang ezyang@meta.com

cc @albanD

[ghstack-poisoned]

pytorch-bot · 2025-08-09T00:55:52Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/160236

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit 3a34b08 with merge base af8c232 ():
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0db2026 Pull-Request: #160236

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 1960626 Pull-Request: #160236

ezyang · 2025-08-10T04:26:11Z

This is ready for review!

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 70c67dd Pull-Request: #160236

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 1c0e378 Pull-Request: #160236

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 6351385 Pull-Request: #160236

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: d70a502 Pull-Request: #160236

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 96c16d3 Pull-Request: #160236

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 36e04aa Pull-Request: #160236

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 5b2acaf Pull-Request: #160236

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 530e4fa Pull-Request: #160236

ezyang · 2025-09-19T13:49:43Z

Comments have been addressed. I think I'll handle csrc removal next PR.

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0dff4de Pull-Request: #160236

ezyang · 2025-09-19T13:54:13Z

@pytorchbot merge

pytorchmergebot · 2025-09-19T13:56:19Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

pytorchmergebot · 2025-09-19T14:01:49Z

Merge failed

Reason: 1 jobs have failed, first few of them are: linux-aarch64 / linux-jammy-aarch64-py3.10 / build

Details for Dev Infra team

Raised by workflow job

[ghstack-poisoned]

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0b9d3d1 Pull-Request: #160236

ezyang · 2025-09-21T02:53:39Z

@pytorchbot merge

pytorchmergebot · 2025-09-21T02:55:26Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0b9d3d1 Pull-Request: #160236

Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: #163340 Approved by: https://github.com/aorenste ghstack dependencies: #160236

The big semantic change (and the reason for this port) is that we no longer monkeypatch Tensor with torchdim's special methods. The new algorithm for handling dispatch is that we first land in `__torch_function__` and we see if a special FCD implementation needs to be dispatch to first, and if there is nothing we fallback to the standard level strategy. Because there is no longer C binding equivalent of classes, we've condensed _C.Dim and Dim together, and similar for Tensor. This resulted in some bugs as the Python API is sometimes different from the C API. I've attempted to disambiguate these but there may still be mistakes (many early bugs were due to this problem). Dim and DimEntry are especially painful as Dim must abide by Tensor equality semantics, but is pointer equality in C (DimEntry doesn't have this problem). Another difference between C/Python that is subtle is we no longer get implicit conversions from Dim to DimEntry, this also caused some bugs. Much of the mechanical porting work was done by claude code. I have a separate PR that deletes functorch._C, but it was useful having dim.cpp to point claude at it so I haven't done it in this PR. From a reviewing perspective, I need to re-review that I didn't forget to port anything, some noticeably missing "small" things are patched_dim_method. I am still in progress of carefully doing a side-by-side review of ports; "simplifications" from claude code were also a major source of bugs. There are two major feature gaps in the implementation: - DelayedTensor and dot handling are not implemented yet. This should be reasonably easy, just need to do it. However, for the purposes of sharded propagation it is actually better not to reconstruct matmuls. - Splitting dimensions with an index like `[x, y]` doesn't work. The problem is that `__getitem__` interprets this as advanced indexing and sends the list to torch.tensor to turn into a tensor, instead of being eligible for `__torch_function__`. I think I might need to hard code a special case for this or something? Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: pytorch#160236 Approved by: https://github.com/zdevito, https://github.com/albanD

Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: pytorch#163340 Approved by: https://github.com/aorenste ghstack dependencies: pytorch#160236

The big semantic change (and the reason for this port) is that we no longer monkeypatch Tensor with torchdim's special methods. The new algorithm for handling dispatch is that we first land in `__torch_function__` and we see if a special FCD implementation needs to be dispatch to first, and if there is nothing we fallback to the standard level strategy. Because there is no longer C binding equivalent of classes, we've condensed _C.Dim and Dim together, and similar for Tensor. This resulted in some bugs as the Python API is sometimes different from the C API. I've attempted to disambiguate these but there may still be mistakes (many early bugs were due to this problem). Dim and DimEntry are especially painful as Dim must abide by Tensor equality semantics, but is pointer equality in C (DimEntry doesn't have this problem). Another difference between C/Python that is subtle is we no longer get implicit conversions from Dim to DimEntry, this also caused some bugs. Much of the mechanical porting work was done by claude code. I have a separate PR that deletes functorch._C, but it was useful having dim.cpp to point claude at it so I haven't done it in this PR. From a reviewing perspective, I need to re-review that I didn't forget to port anything, some noticeably missing "small" things are patched_dim_method. I am still in progress of carefully doing a side-by-side review of ports; "simplifications" from claude code were also a major source of bugs. There are two major feature gaps in the implementation: - DelayedTensor and dot handling are not implemented yet. This should be reasonably easy, just need to do it. However, for the purposes of sharded propagation it is actually better not to reconstruct matmuls. - Splitting dimensions with an index like `[x, y]` doesn't work. The problem is that `__getitem__` interprets this as advanced indexing and sends the list to torch.tensor to turn into a tensor, instead of being eligible for `__torch_function__`. I think I might need to hard code a special case for this or something? Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: pytorch#160236 Approved by: https://github.com/zdevito, https://github.com/albanD

Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: pytorch#163340 Approved by: https://github.com/aorenste ghstack dependencies: pytorch#160236

The big semantic change (and the reason for this port) is that we no longer monkeypatch Tensor with torchdim's special methods. The new algorithm for handling dispatch is that we first land in `__torch_function__` and we see if a special FCD implementation needs to be dispatch to first, and if there is nothing we fallback to the standard level strategy. Because there is no longer C binding equivalent of classes, we've condensed _C.Dim and Dim together, and similar for Tensor. This resulted in some bugs as the Python API is sometimes different from the C API. I've attempted to disambiguate these but there may still be mistakes (many early bugs were due to this problem). Dim and DimEntry are especially painful as Dim must abide by Tensor equality semantics, but is pointer equality in C (DimEntry doesn't have this problem). Another difference between C/Python that is subtle is we no longer get implicit conversions from Dim to DimEntry, this also caused some bugs. Much of the mechanical porting work was done by claude code. I have a separate PR that deletes functorch._C, but it was useful having dim.cpp to point claude at it so I haven't done it in this PR. From a reviewing perspective, I need to re-review that I didn't forget to port anything, some noticeably missing "small" things are patched_dim_method. I am still in progress of carefully doing a side-by-side review of ports; "simplifications" from claude code were also a major source of bugs. There are two major feature gaps in the implementation: - DelayedTensor and dot handling are not implemented yet. This should be reasonably easy, just need to do it. However, for the purposes of sharded propagation it is actually better not to reconstruct matmuls. - Splitting dimensions with an index like `[x, y]` doesn't work. The problem is that `__getitem__` interprets this as advanced indexing and sends the list to torch.tensor to turn into a tensor, instead of being eligible for `__torch_function__`. I think I might need to hard code a special case for this or something? Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: pytorch#160236 Approved by: https://github.com/zdevito, https://github.com/albanD

Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: pytorch#163340 Approved by: https://github.com/aorenste ghstack dependencies: pytorch#160236

Update

7bf28fa

[ghstack-poisoned]

github-actions bot requested review from SherlockNoMad, albanD, antoniojkim, bdhirsh and miladm August 9, 2025 00:56

ezyang added a commit that referenced this pull request Aug 9, 2025

[WIP] torchdim Python port

c308150

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0db2026 Pull-Request: #160236

ezyang mentioned this pull request Aug 9, 2025

Delete Python reference implementation from torchdim, as it is untested #160115

Closed

Update

333fd1e

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Aug 10, 2025

[WIP] torchdim Python port

b2db630

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 1960626 Pull-Request: #160236

ezyang mentioned this pull request Aug 10, 2025

Detect torch function in lists as well #160256

Closed

ezyang changed the title ~~[WIP] torchdim Python port~~ torchdim Python port Aug 10, 2025

ezyang requested a review from zdevito August 10, 2025 04:18

ezyang mentioned this pull request Aug 10, 2025

State of Torch Named Tensors #60832

Open

Update

172e55e

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Aug 10, 2025

[WIP] torchdim Python port

59f0a79

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 70c67dd Pull-Request: #160236

Update

afaf3dc

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Aug 12, 2025

[WIP] torchdim Python port

5a8498e

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 1c0e378 Pull-Request: #160236

Update

5ed2137

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Aug 12, 2025

[WIP] torchdim Python port

b82d3c2

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 6351385 Pull-Request: #160236

Update

6c91a8d

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Aug 12, 2025

[WIP] torchdim Python port

eef2aab

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: d70a502 Pull-Request: #160236

Update

a49ee63

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Aug 12, 2025

[WIP] torchdim Python port

5194874

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 96c16d3 Pull-Request: #160236

ezyang added the suppress-bc-linter Suppresses the failures of API backward-compatibility linter (Lint/bc_linter) label Aug 12, 2025

Update

c77ccf1

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Aug 12, 2025

[WIP] torchdim Python port

a22896c

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 36e04aa Pull-Request: #160236

ezyang added the skip-pr-sanity-checks label Aug 13, 2025

ezyang added a commit that referenced this pull request Sep 18, 2025

[WIP] torchdim Python port

69131b4

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 5b2acaf Pull-Request: #160236

Update

aa90165

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Sep 19, 2025

[WIP] torchdim Python port

e25b5bc

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 530e4fa Pull-Request: #160236

Update

7c42487

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Sep 19, 2025

[WIP] torchdim Python port

9292e9e

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0dff4de Pull-Request: #160236

ezyang mentioned this pull request Sep 19, 2025

Delete functorch C extension entirely. #163340

Closed

pytorchmergebot added the merging label Sep 19, 2025

pytorchmergebot removed the merging label Sep 19, 2025

Update

3a34b08

[ghstack-poisoned]

ezyang added a commit that referenced this pull request Sep 20, 2025

[WIP] torchdim Python port

798a3e0

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0b9d3d1 Pull-Request: #160236

pytorchmergebot added the merging label Sep 21, 2025

pytorchmergebot added the Merged label Sep 21, 2025

pytorchmergebot closed this in 97eb7a2 Sep 21, 2025

pytorchmergebot removed the merging label Sep 21, 2025

pytorchmergebot pushed a commit that referenced this pull request Sep 21, 2025

[WIP] torchdim Python port

c15d9b6

Signed-off-by: Edward Yang <ezyang@meta.com> ghstack-source-id: 0b9d3d1 Pull-Request: #160236

pytorchmergebot pushed a commit that referenced this pull request Sep 21, 2025

Delete functorch C extension entirely. (#163340)

1faf636

Signed-off-by: Edward Yang <ezyang@meta.com> Pull Request resolved: #163340 Approved by: https://github.com/aorenste ghstack dependencies: #160236

github-actions bot deleted the gh/ezyang/3127/head branch October 22, 2025 02:15

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

torchdim Python port#160236

torchdim Python port#160236
ezyang wants to merge 16 commits intogh/ezyang/3127/basefrom
gh/ezyang/3127/head

ezyang commented Aug 9, 2025 •

edited

Loading

Uh oh!

pytorch-bot bot commented Aug 9, 2025 •

edited

Loading

Uh oh!

ezyang commented Aug 10, 2025

Uh oh!

ezyang commented Sep 19, 2025

Uh oh!

ezyang commented Sep 19, 2025

Uh oh!

pytorchmergebot commented Sep 19, 2025

Uh oh!

pytorchmergebot commented Sep 19, 2025

Uh oh!

ezyang commented Sep 21, 2025

Uh oh!

pytorchmergebot commented Sep 21, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

Conversation

ezyang commented Aug 9, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

pytorch-bot bot commented Aug 9, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/160236

✅ No Failures

Uh oh!

ezyang commented Aug 10, 2025

Uh oh!

ezyang commented Sep 19, 2025

Uh oh!

ezyang commented Sep 19, 2025

Uh oh!

pytorchmergebot commented Sep 19, 2025

Merge started

Uh oh!

pytorchmergebot commented Sep 19, 2025

Merge failed

Uh oh!

ezyang commented Sep 21, 2025

Uh oh!

pytorchmergebot commented Sep 21, 2025

Merge started

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

ezyang commented Aug 9, 2025 •

edited

Loading

pytorch-bot bot commented Aug 9, 2025 •

edited

Loading