Skip to content

fix: resolve batched service header columns - #2729

Merged
Rana Singh (ranadeepsingh) merged 2 commits into
microsoft:masterfrom
leninworld:fix-batched-service-header-cols-issue-2064
Sep 21, 2026
Merged

Rana Singh (ranadeepsingh) merged 2 commits into
microsoft:masterfrom
leninworld:fix-batched-service-header-cols-issue-2064

Conversation

@leninworld

@leninworld Lenin Mookiah (leninworld) commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Related Issues/PRs

Fixes #2064

What changes are proposed in this pull request?

Automatically batched Cognitive Service transformers now resolve request-header service parameters from the batched values instead of casting the entire Spark array to a scalar.

  • Resolve string credentials (subscription key, AAD token, and custom authorization header) from the first non-blank value in a batch.
  • Resolve custom and telemetry header maps from the first non-empty sanitized map.
  • Apply the same batch-aware path to the Fabric fallback credential check.
  • Keep payload parameters such as text and language on the existing typed path so document-aligned arrays remain intact.
  • Raise a parameter-specific error for incompatible present header values, without echoing their contents. Empty arrays/maps are treated as absent; this is not full Spark column-schema validation.
  • Document that one credential authenticates each HTTP batch; users who require per-row credential selection should set batch size to 1.

Batching behavior is otherwise unchanged. No warning is emitted for multiple distinct credentials within one batch.

This is a shared-header fix, not a general solution for every column-bound request parameter. Independent public-API probes on the current packaged code still reproduce automatic-batching failures for modelVersionCol and loggingOptOutCol in AnalyzeText and AnalyzeTextLongRunningOperations. The legacy TextSentiment.modelVersionCol path sends the batched collection's string representation as the query value. Those request-body/query paths are unchanged by this PR and need a separate fix. The credential-column control passes in all three stages.

Translator's separate subscriptionRegion parameter is unchanged. Its scalar region column is not automatically batched; manually supplied array-valued regions remain unsupported.

How is this patch tested?

  • I have written tests (not required for typo or doc fix) and confirmed the proposed feature/bug-fix/change works.

Local validation with JDK 11 and Spark 3.5.0:

  • cognitive/Test/compile passed.
  • Cognitive production and test scalastyle passed with zero errors and warnings.
  • CognitiveServiceBaseSuite: 12/12 passed, covering batched string/map headers, null and blank credentials, Fabric fallback, invalid element types, public batching/flattening, and payload isolation.

Docker integration used Spark 3.5.1, Java 11.0.22, and Linux x86-64. A real TextSentiment stage loaded the branch-built Cognitive JAR and called a credential-free local mock endpoint. Two rows were combined into one request, the first usable subscription key authenticated the batch, and FlattenBatch preserved both original row keys.

Validation evidence

Local compilation, style, and focused Scala regression suite:

issue-2064-local-validation

Docker Spark end-to-end regression using the branch-built artifact:

issue-2064-docker-spark

Repaired TextSentiment function loaded and transformed the automatically batched subscriptionKeyCol:

issue-2064-api-loaded

Does this PR change any dependencies?

  • No. You can skip this section.

Does this PR add a new feature? If so, have you added samples on website?

  • No. You can skip this section.

Maintainer follow-up

Added 06f84e3a4a on the original contributor branch without rewriting the contributor commit.

  • Fixed a cross-version gap: Seq means immutable sequences in Scala 2.13. Matching scala.collection.Seq also accepts Spark's mutable arrays and keeps the existing lazy traversal without copying a batch.
  • Added five credential-free loopback HTTP tests through real TextSentiment and AnalyzeHealthText transforms. They cover automatic and manual batching, partial batches, per-row routing with batch size one, copy/save/load, asynchronous polling credentials, text/language alignment, and output keys.
  • Added four helper regressions for mutable arrays, auth precedence with lazy Fabric fallback, blank credentials, and nullable header maps. Strengthened invalid-type redaction and made the existing batch test deterministic.
  • Clarified that batching does not group by credential and later header maps are not merged.

Independent local evidence for the final source:

Runtime Targeted Scala results
Spark 3.5.0 / Scala 2.12.17 / JDK 11 22 passed, no failures or skips
Spark 4.0.1 / Scala 2.13.16 / JDK 17 22 passed, no failures or skips
Spark 4.1.1 / Scala 2.13.17 / JDK 17 22 passed, no failures or skips

The Spark 4 runs replayed the patch on current port heads 7251246d45 and b4ca894139, preserving their collection normalization and runtime settings. No shared port branch was modified.

The public regression suite was also run against the pre-fix master implementation. Four automatic-batching tests reproduced the reported WrappedArray-to-String exception; the scalar manual-batch control passed. Replaying the contributor helper on both Spark 4 ports exposed mutable-array rejection before the maintainer correction.

Both generated Python stages passed the loopback regression using the packaged Core/Cognitive JARs, with class-loading provenance checked. This local Python probe used Python 3.12; CI provides the repository-pinned Python environment. Core/Cognitive compile, test compile, Scala style, code generation, pinned Black, and notebook JSON checks passed.

Current-head Azure validation passed after /azp run: build 236902583, with trigger reason pullRequest and merge source 058aa19fa3164b0f8336871adc621d7e96cd90f6. Its parents are master 3878cae184909569328ea41a6aaccf06f9e80d36 and PR head 06f84e3a4a228f05217fa4e43be20587b4de2bf0.

  • All 64 Azure jobs succeeded. Published results contain 3,610 passed tests, no failures, and 22 pre-existing explicitly ignored/skipped tests. Every test run's build ID was checked.
  • All 22 affected tests passed in Azure, including all nine maintainer additions; none of these were skipped.
  • Current-head GitHub workflows succeeded. Website deployment itself is intentionally skipped for PRs; its build passed.
  • CI's Spark 4.1 replay is a compile/test-compile check, not a runtime test run. The 22-test Spark 4.0 and 4.1 runtime evidence above comes from the independent local port replays.
  • Non-gating warnings were inspected: WebsiteSamplesTests had no Scala coverage file to publish, the compile-only Spark 4.1 replay had no executed-test XML to publish, and the Docker agent reported low disk space. These jobs completed successfully; the affected test results were present.
  • SynapseML-Internal compatibility is excluded for fork PRs. This run is not evidence of live Internal AI Functions validation.

The final readiness audit found the PR zero commits behind master, no failed/pending/missing required checks, no unresolved review threads, and no suppressed current-head findings. The current-head automated review reported no findings; its scope observations are addressed in this response. CLA status passed. Contributor sign-off on the maintainer additions and human maintainer approval remain outstanding; this PR has not been merged.

The earlier build, 236892351, had published 3,592 passing tests and 22 not-executed tests at an intermediate inspection, but hit an Azure certificate-name mismatch during coverage publication and was subsequently canceled after the follow-up push. Those intermediate counts are not a clean or final CI result.

No workflow, dependency, public JVM signature, serialized parameter shape, or payload-handling changes are included.

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
There may be pipelines that require an authorized user to comment /azp run to run.

@github-actions

Copy link
Copy Markdown

Hey Lenin Mookiah (@leninworld) 👋!
Thank you so much for contributing to our repository 🙌.
Someone from SynapseML Team will be reviewing this pull request soon.

We use semantic commit messages to streamline the release process.
Before your pull request can be merged, you should make sure your first commit and PR title start with a semantic prefix.
This helps us to create release messages and credit you for your hard work!

Examples of commit messages with semantic prefixes:

  • fix: Fix LightGBM crashes with empty partitions
  • feat: Make HTTP on Spark back-offs configurable
  • docs: Update Spark Serving usage
  • build: Add codecov support
  • perf: improve LightGBM memory usage
  • refactor: make python code generation rely on classes
  • style: Remove nulls from CNTKModel
  • test: Add test coverage for CNTKModel

To test your commit locally, please follow our guild on building from source.
Check out the developer guide for additional guidance on testing your change.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

No unresolved review issues remain, and all assessments indicate approval readiness.

Review effort: Lite
Findings: None

What changed in this PR

Fixes batched Cognitive Service authentication by resolving credentials and headers from batched values while preserving payload arrays.

Changes:

  • Adds batch-aware credential and header-map resolution.
  • Updates Fabric fallback handling and validates incompatible types.
  • Adds regression tests and documents batch credential behavior.
File Description
docs/​Explore Algorithms/​AI Services/​Advanced Usage - Async, Batching, and Multi-Key.ipynb Documents per-batch credential behavior.
cognitive/​src/​test/​scala/​com/​microsoft/​azure/​synapse/​ml/​services/​CognitiveServiceBaseSuite.scala Tests batching, fallback, validation, and payload preservation.
cognitive/​src/​main/​scala/​com/​microsoft/​azure/​synapse/​ml/​services/​ServiceHeaderValues.scala Resolves and validates batched header values.
cognitive/​src/​main/​scala/​com/​microsoft/​azure/​synapse/​ml/​services/​CognitiveServiceBase.scala Applies batch-aware authentication and header resolution.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@ranadeepsingh

Copy link
Copy Markdown
Collaborator

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 1 pipeline(s).

## Summary
Keep the contributor's header-only batch resolution and accept mutable as
well as immutable Scala sequences. Add credential-free public-transformer
regressions and clarify per-batch header selection.

## Prompting Intent
Review and repair the submitted PR on its existing contributor branch.
Prove that it fixes the reported issue, check security and regressions, and
complete current-head review and CI rather than stopping at a queued build.

## Linked Sources
- Submitted PR: microsoft#2729
- Original issue and reproduction: microsoft#2064
- Contributor implementation: microsoft@14c59c4

## Rationale
Scala 2.13's unqualified Seq excludes mutable Spark array representations.
Matching scala.collection.Seq retains lazy, allocation-light traversal
without changing public signatures, parameter serialization, or payload
handling. Replaying the contributor implementation on both Spark ports
demonstrated the rejection before this adjustment.

Loopback HTTP tests exercise TextSentiment and AnalyzeHealthText through
automatic batching, request creation, asynchronous polling, response parsing,
and flattening. They also check partial batches, scalar manual batches,
batch-size-one routing, copy/save/load, document languages, and output alignment.
The original master implementation fails the four automatic-batching cases
with the issue's WrappedArray-to-String cast, while the scalar control passes.

Strengthen helper coverage for auth precedence, lazy Fabric fallback, nullable
maps, invalid-type redaction, and mutable arrays. Make the batching helper
test deterministic with one partition. Document that batches do not group by
credential and header maps from later rows are not merged.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot AI review requested due to automatic review settings September 21, 2026 10:13
@ranadeepsingh

Copy link
Copy Markdown
Collaborator

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines:
Successfully started running 1 pipeline(s).

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Batched subscription regions remain unsupported, and incompatible empty arrays/maps bypass schema type validation.

Review effort: Lite
Findings: None

@ranadeepsingh

Copy link
Copy Markdown
Collaborator

I checked both observations in the current-head review.

Subscription regions. subscriptionRegion belongs to Translator's separate request path, not the Text Analytics automatic-batching path from #2064. Translator's customGetInternalTransformer reshapes only the explicitly named payload columns, such as text and toLanguage; it does not batch whole rows or turn a scalar region column into an array. That code is unchanged in this PR. Manually supplying an array-valued region remains unsupported, and this PR does not claim to add it.

Empty values and schema validation. The new helper validates values it consumes, not every declared Spark column type. Empty arrays and empty maps contain no usable credential/header, so they are treated as absent. Present incompatible values raise parameter-specific errors without echoing their contents. Full schema validation, including rejection of empty arrays with incompatible declared element types, is not part of this fix. I clarified the PR description to avoid claiming otherwise.

Neither observation invalidates the reproduced #2064 fix or identifies a regression in the supported path. The distinction matters: the four automatic-batching public regressions fail on the original master code and pass with this patch, and all 22 targeted tests pass on each of Spark 3.5, 4.0, and 4.1. No review threads or suppressed findings were present on 06f84e3a4a.

@ranadeepsingh

Copy link
Copy Markdown
Collaborator

Thanks Lenin Mookiah (@leninworld) for fixing the batched credential-column crash in #2064.

I kept your shared-header approach and pushed 06f84e3a4a to this PR without rewriting your commit. The production adjustment is to match scala.collection.Seq, since unqualified Seq excludes Spark's mutable arrays under Scala 2.13. The follow-up also adds five public-transformer HTTP regressions, four helper regressions, and documentation clarifying first-usable credentials per batch, independent header-column selection, and no credential grouping or map merging.

The public regressions reproduced four automatic-batching failures on pre-fix master and passed with the patch. All 22 targeted Scala tests passed locally on Spark 3.5, 4.0, and 4.1. The generated Python APIs also passed the credential-free loopback checks against the packaged JARs.

Azure build 236902583, triggered with /azp run, passed for the merge of head 06f84e3a4a into current master:

  • All 64 jobs succeeded.
  • Published results show 3,610 passed tests, zero failures, and 22 pre-existing explicitly ignored/skipped tests.
  • All 22 affected tests ran and passed, including all nine additions.
  • Current-head GitHub workflows passed. Website deployment is intentionally PR-skipped; its build passed.

The final review audit found no unresolved threads or suppressed current-head findings. I addressed the automated review's scope observations here. The PR description records the inspected CI warnings and distinguishes CI's Spark 4.1 compile replay from the local runtime tests.

The description also makes the scope explicit: this fixes shared headers, not all column-bound request settings. Separate probes still reproduce pre-existing modelVersionCol and loggingOptOutCol batching problems in unchanged request-body/query paths. Those need a separate follow-up rather than expanding your contribution into a batching redesign.

Please review my additions and sign off if they fit your intent. If they do not, feel free to revert my commit, or let me know and I will revert it.

CLA status is green. Your sign-off on these additions and human maintainer approval are still outstanding. I have not merged the PR.

@leninworld

Copy link
Copy Markdown
Contributor Author

Thanks Lenin Mookiah (Lenin Mookiah (@leninworld)) for fixing the batched credential-column crash in #2064.

I kept your shared-header approach and pushed 06f84e3a4a to this PR without rewriting your commit. The production adjustment is to match scala.collection.Seq, since unqualified Seq excludes Spark's mutable arrays under Scala 2.13. The follow-up also adds five public-transformer HTTP regressions, four helper regressions, and documentation clarifying first-usable credentials per batch, independent header-column selection, and no credential grouping or map merging.

The public regressions reproduced four automatic-batching failures on pre-fix master and passed with the patch. All 22 targeted Scala tests passed locally on Spark 3.5, 4.0, and 4.1. The generated Python APIs also passed the credential-free loopback checks against the packaged JARs.

Azure build 236902583, triggered with /azp run, passed for the merge of head 06f84e3a4a into current master:

  • All 64 jobs succeeded.
  • Published results show 3,610 passed tests, zero failures, and 22 pre-existing explicitly ignored/skipped tests.
  • All 22 affected tests ran and passed, including all nine additions.
  • Current-head GitHub workflows passed. Website deployment is intentionally PR-skipped; its build passed.

The final review audit found no unresolved threads or suppressed current-head findings. I addressed the automated review's scope observations here. The PR description records the inspected CI warnings and distinguishes CI's Spark 4.1 compile replay from the local runtime tests.

The description also makes the scope explicit: this fixes shared headers, not all column-bound request settings. Separate probes still reproduce pre-existing modelVersionCol and loggingOptOutCol batching problems in unchanged request-body/query paths. Those need a separate follow-up rather than expanding your contribution into a batching redesign.

Please review my additions and sign off if they fit your intent. If they do not, feel free to revert my commit, or let me know and I will revert it.

CLA status is green. Your sign-off on these additions and human maintainer approval are still outstanding. I have not merged the PR.

Thanks for the detailed follow-up. I reviewed commit 06f84e3a4a and the updated PR. The scala.collection.Seq adjustment preserves the shared-header approach while supporting the mutable Spark collection representations, and the added regression tests and documentation match the intended scope of #2064.

I also agree with keeping the other column-bound settings and full schema validation outside this PR. The additions fit my intent, and I’m signing off on the current head. I have no further changes to request.

@ranadeepsingh
Rana Singh (ranadeepsingh) merged commit 714d365 into microsoft:master Sep 21, 2026
81 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] Using subscriptionKeyCol argument results in SparkException [FAILED_EXECUTE_UDF]

4 participants