DCP HF reader: use safe_open instead of reading the bytes by ankitageorge · Pull Request #159406 · pytorch/pytorch

ankitageorge · 2025-07-29T19:46:54Z

Stack from ghstack (oldest at bottom):

Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after.

Differential Revision: D78994259

cc @H-Huang @awgu @wanchaol @fegin @fduwjj @wz337 @wconstab @d4l3k @pragupta

Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/) [ghstack-poisoned]

pytorch-bot · 2025-07-29T19:46:57Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/159406

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki or our office hours

Note: Links to docs will display an error until the docs builds have been completed.

⏳ 1 Pending, 1 Unrelated Failure

As of commit 12b6ef4 with merge base 93da995 ():

UNSTABLE - The following job is marked as unstable, possibly due to flakiness on trunk:

pull / linux-jammy-py3_9-clang9-xla / test (xla, 1, 1, lf.linux.12xlarge, unstable) (gh) (#158876)
/var/lib/jenkins/workspace/xla/torch_xla/csrc/runtime/BUILD:476:14: Compiling torch_xla/csrc/runtime/xla_util_test.cpp failed: (Exit 1): gcc failed: error executing CppCompile command (from target //torch_xla/csrc/runtime:xla_util_test) /usr/bin/gcc -U_FORTIFY_SOURCE -fstack-protector -Wall -Wunused-but-set-parameter -Wno-free-nonheap-object -fno-omit-frame-pointer -g0 -O2 '-D_FORTIFY_SOURCE=1' -DNDEBUG -ffunction-sections ... (remaining 229 arguments skipped)

This comment was automatically generated by Dr. CI and updates every 15 minutes.

Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/) ghstack-source-id: 298671089 Pull Request resolved: #159406

facebook-github-bot · 2025-07-29T19:47:16Z

This pull request was exported from Phabricator. Differential Revision: D78994259

Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/) cc H-Huang awgu wanchaol fegin fduwjj wz337 wconstab d4l3k pragupta [ghstack-poisoned]

Pull Request resolved: #159406 Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/) ghstack-source-id: 300113574

facebook-github-bot · 2025-08-01T16:38:10Z

This pull request was exported from Phabricator. Differential Revision: D78994259

Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/) cc H-Huang awgu wanchaol fegin fduwjj wz337 wconstab d4l3k pragupta [ghstack-poisoned]

Pull Request resolved: #159406 Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. ghstack-source-id: 300151112 Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/)

facebook-github-bot · 2025-08-01T17:30:59Z

This pull request was exported from Phabricator. Differential Revision: D78994259

saumishr

LGTM!

Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/) cc H-Huang awgu wanchaol fegin fduwjj wz337 wconstab d4l3k pragupta [ghstack-poisoned]

facebook-github-bot · 2025-08-07T13:49:09Z

This pull request was exported from Phabricator. Differential Revision: D78994259

facebook-github-bot · 2025-08-07T17:20:28Z

@pytorchbot merge

(Initiating merge automatically since Phabricator Diff has merged)

pytorchmergebot · 2025-08-07T17:22:25Z

Starting merge as part of PR stack under #159681

pytorchmergebot · 2025-08-07T17:22:36Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

Get rid of the logic to read the metadata from the header of the safetensors file manually and use the functions as part of safe_open() to get the metadata. This is much cleaner and allows us to not rely on our own custom methods to get metadata, but use safetensors provided APIs Differential Revision: [D79460272](https://our.internmc.facebook.com/intern/diff/D79460272/) Pull Request resolved: #159681 Approved by: https://github.com/saumishr ghstack dependencies: #159405, #159406

…9406) Reading the bytes and converting to tensors is much slower than using safe_open. For a 8B model across 8 ranks, took ~30s to load before this change and ~4s after. Differential Revision: [D78994259](https://our.internmc.facebook.com/intern/diff/D78994259/) Pull Request resolved: pytorch#159406 Approved by: https://github.com/saumishr ghstack dependencies: pytorch#159405

Get rid of the logic to read the metadata from the header of the safetensors file manually and use the functions as part of safe_open() to get the metadata. This is much cleaner and allows us to not rely on our own custom methods to get metadata, but use safetensors provided APIs Differential Revision: [D79460272](https://our.internmc.facebook.com/intern/diff/D79460272/) Pull Request resolved: pytorch#159681 Approved by: https://github.com/saumishr ghstack dependencies: pytorch#159405, pytorch#159406

ankitageorge mentioned this pull request Jul 29, 2025

HF component update to not use fsspec components #159405

Closed

pytorch-bot bot added oncall: distributed Add this issue/PR to distributed oncall triage queue release notes: distributed (checkpoint) labels Jul 29, 2025

facebook-github-bot added the fb-exported label Jul 29, 2025

saumishr reviewed Aug 1, 2025

View reviewed changes

saumishr approved these changes Aug 1, 2025

View reviewed changes

pytorch-bot bot added the ciflow/trunk Trigger trunk jobs on your pull request label Aug 1, 2025

ankitageorge mentioned this pull request Aug 1, 2025

Use only safetensors APIs in HFStorageReader #159681

Closed

pytorchmergebot added the merging label Aug 7, 2025

pytorchmergebot closed this in 0b187b3 Aug 7, 2025

pytorchmergebot added Merged and removed merging labels Aug 7, 2025

github-actions bot deleted the gh/ankitageorge/19/head branch September 7, 2025 02:14

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

DCP HF reader: use safe_open instead of reading the bytes#159406

DCP HF reader: use safe_open instead of reading the bytes#159406
ankitageorge wants to merge 4 commits intogh/ankitageorge/19/basefrom
gh/ankitageorge/19/head

ankitageorge commented Jul 29, 2025 •

edited

Loading

Uh oh!

pytorch-bot bot commented Jul 29, 2025 •

edited

Loading

Uh oh!

facebook-github-bot commented Jul 29, 2025

Uh oh!

facebook-github-bot commented Aug 1, 2025

Uh oh!

facebook-github-bot commented Aug 1, 2025

Uh oh!

saumishr left a comment

Uh oh!

facebook-github-bot commented Aug 7, 2025

Uh oh!

facebook-github-bot commented Aug 7, 2025

Uh oh!

pytorchmergebot commented Aug 7, 2025

Uh oh!

pytorchmergebot commented Aug 7, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

Conversation

ankitageorge commented Jul 29, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

pytorch-bot bot commented Jul 29, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/159406

⏳ 1 Pending, 1 Unrelated Failure

Uh oh!

facebook-github-bot commented Jul 29, 2025

Uh oh!

facebook-github-bot commented Aug 1, 2025

Uh oh!

facebook-github-bot commented Aug 1, 2025

Uh oh!

saumishr left a comment

Choose a reason for hiding this comment

Uh oh!

facebook-github-bot commented Aug 7, 2025

Uh oh!

facebook-github-bot commented Aug 7, 2025

Uh oh!

pytorchmergebot commented Aug 7, 2025

Uh oh!

pytorchmergebot commented Aug 7, 2025

Merge started

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

ankitageorge commented Jul 29, 2025 •

edited

Loading

pytorch-bot bot commented Jul 29, 2025 •

edited

Loading