Skip to content

[Fix][Relax][Frontend][Torch] Apply the scale and input_scale arguments of aten.elu - #20515

Open
Arthur031221 wants to merge 2 commits into
apache:mainfrom
Arthur031221:fix-elu-scale-main
Open

Arthur031221 wants to merge 2 commits into
apache:mainfrom
Arthur031221:fix-elu-scale-main

Conversation

@Arthur031221

Copy link
Copy Markdown

Anyone importing a model that uses torch.nn.SELU or F.selu through from_exported_program gets outputs that are about 5% too small, because the importer drops the SELU scale.

from_exported_program runs run_decompositions() by default, which rewrites aten.selu to aten.elu(x, alpha, scale). The elu.default converter only reads alpha, so scale (1.0507...) and input_scale are ignored. The expected IR in test_extended_unary_ops for SELU had the same omission, so the test passed on the wrong result.

Reproduction on F.selu(torch.tensor([[-2.0, -0.5, 0.0, 1.5]])), compiled with llvm and run on the VM:

before
torch [[-1.5201665  -0.69175816  0.          1.5760515 ]]
tvm   [[-1.4468117  -0.65837777  0.          1.5       ]]
after
torch [[-1.5201665  -0.69175816  0.          1.5760515 ]]
tvm   [[-1.5201665  -0.6917582   0.          1.5760515 ]]

Changes:

  • ExportedProgramImporter._elu now applies scale and input_scale from aten.elu. With the defaults (1 and 1) it still calls the existing converter, so plain F.elu and nn.ELU produce the same IR as before. The fx importer is untouched.
  • The SELU expectation in test_extended_unary_ops gets the final multiply by the SELU scale.
  • New test_selu_and_elu_scale_arguments compares against PyTorch numerically for F.selu and for aten.elu(x, 0.5, 2.0, 3.0). It fails without the converter change (3 of 4 elements mismatch, max absolute difference 0.076) and passes with it.

Tests: tests/python/relax/test_frontend_from_exported_program.py and test_frontend_from_fx.py give 437 passed, 1 failed, 3 skipped with the change; the one failure is test_norm, which fails the same way before it.

@tlopex tlopex left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This ReLU formulation only works when input_scale >= 0. For example, with x=1.5, alpha=0.5, scale=2, and input_scale=-1, PyTorch returns 3.0, but this implementation returns approximately 2.22313. Could we select the branch based on the original input (where(x < 0, alpha * (exp(input_scale * x) - 1), x)), then apply scale?

And could you move the test to the end of the file and make it a structural check instead of numerical check?

@Arthur031221

Copy link
Copy Markdown
Author

Thanks, you are right about negative input_scale. I changed the branch test to use the original input, then apply scale after selecting the branch. I also moved the added test to the end of the file and made it a structural check.

For aten.elu([-2, -0.5, 0, 1.5], 0.5, 2, -1), the compiled result is [6.3890562, 0.6487212, 0, 3]; PyTorch gives [6.3890562, 0.6487213, 0, 3]. The maximum absolute difference is 5.9604645e-8. The new structural test fails with the previous importer and passes with this change. The existing SELU structural test also passes with its expectation updated.

The PR description refers to the numerical test from the earlier revision. This follow-up replaces that test with the structural check you requested.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants