feat: abstraction of xla::OpSharding proto using wrapper class #9467

kvshbg-aws · 2025-07-10T00:29:51Z

This PR includes the changes related to abstracting xla::OpSharidng proto object into a torch_xla::OpSharding wrapper class.

This new class object will not have the requirements of xla::OpSharding (however, it will be an extension xla::OpSharding proto defined over here).
We have defined the wrapper class in torch/xla which will construct an xla::OpSharding object with additional fields such as global_device_ids/global_tile_assignment and will have forwarded/proxy functions to xla::OpSharding . These forwarded functions will help user still make use of the same xla::OpSharding APIs as they normally would. We can also define torch_xla specific functions in this wrapper class to further use the extra fields that were stored during the initialization of the OpSharding object. This approach also allows the flexibility of converting the torch_xla::OpSharding object back to xla::OpSharding while lowering into HLO, thus, giving user the flexibility to use the abstracted class (and other additional fields stored) anywhere in the code base as needed, this is particularly useful since the XLA's HLOs are 0th indexed, hence we need to use the normalized_device_ids (starting from index 0) when lowering the program into the HLO, whereas we can still use the denormalized/global_device_ids in other places such as inside pjrt client to set the device_assignment using the user specified device_ids.

Component diagram for reference -

Ref issue - #9390

…_assignment() is empty

…o_data

…ment

kvshbg-aws force-pushed the kvshbg-aws/local-spmd-abstraction branch from d0502ab to 7fc15ea Compare July 10, 2025 18:34

qihqi requested review from rpsilva-aws and pgmoka July 11, 2025 04:20

kvshbg-aws force-pushed the kvshbg-aws/local-spmd-abstraction branch 4 times, most recently from 7c4a3cd to 1d55ae9 Compare July 16, 2025 23:57

kvshbg-aws added 7 commits July 23, 2025 21:20

feat: abstraction of xla::OpSharding proto using wrapper class

ca1c72f

fix for failing ci/cd tests

0059e30

small fix to remove unwanted var

320d0f1

bug fix: set value of denormalized_tile_assignment when sharding.tile…

5bc3fb6

…_assignment() is empty

use tensors to get denormalized_tile_assignment directly instead of p…

2e34d5a

…o_data

use tensors instead of paramters_data to get denormalized_tile_assign…

b390dd5

…ment

linter fix

1ddbb1b

kvshbg-aws force-pushed the kvshbg-aws/local-spmd-abstraction branch from 2756c1a to 1ddbb1b Compare July 23, 2025 21:20

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

feat: abstraction of xla::OpSharding proto using wrapper class #9467

feat: abstraction of xla::OpSharding proto using wrapper class #9467

Uh oh!

kvshbg-aws commented Jul 10, 2025

Uh oh!

Uh oh!

feat: abstraction of xla::OpSharding proto using wrapper class #9467

Are you sure you want to change the base?

feat: abstraction of xla::OpSharding proto using wrapper class #9467

Uh oh!

Conversation

kvshbg-aws commented Jul 10, 2025

Uh oh!

Uh oh!