Skip to content

chore: support Spark 4.2#2289

Merged
linhr merged 14 commits into
mainfrom
spark-4-2
Jul 27, 2026
Merged

chore: support Spark 4.2#2289
linhr merged 14 commits into
mainfrom
spark-4-2

Conversation

@linhr

@linhr linhr commented Jul 25, 2026

Copy link
Copy Markdown
Contributor

No description provided.

@github-actions

github-actions Bot commented Jul 25, 2026

Copy link
Copy Markdown

Gold Data Report

Notes
  1. The tables below show the number of true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN) in gold data input processing.
  2. A positive input is a valid test case, while a negative input is a test case that is expected to fail.

Commit Information

Commit Revision Branch
After 57d3fb7 refs/pull/2289/merge
Before 6d1e068 main

Summary

Commit TP TN FP FN Total
After 2222 217 50 313 2802
Before 2197 195 51 180 2623

Details

Gold Data Metrics
Group File Commit TP TN FP FN Total
spark data_type.json After 50 11 2 3 66
Before 48 11 2 3 64
expression/case.json After 5 0 0 0 5
Before 5 0 0 0 5
expression/cast.json After 4 0 0 0 4
Before 4 0 0 0 4
expression/current.json After 3 0 0 0 3
Before 3 0 0 0 3
expression/date.json After 4 0 1 0 5
Before 4 0 1 0 5
expression/interval.json After 346 4 1 0 351
Before 346 4 1 0 351
expression/large.json After 2 0 0 0 2
Before 2 0 0 0 2
expression/like.json After 29 10 0 0 39
Before 29 10 0 0 39
expression/misc.json After 111 5 1 1 118
Before 111 5 1 1 118
expression/numeric.json After 31 6 1 0 38
Before 31 6 1 0 38
expression/string.json After 18 1 0 0 19
Before 18 1 0 0 19
expression/timestamp.json After 7 0 3 0 10
Before 7 0 3 0 10
expression/window.json After 73 0 1 0 74
Before 73 0 1 0 74
function/agg.json After 167 0 0 31 198
Before 165 0 0 21 186
function/array.json After 44 0 0 0 44
Before 44 0 0 0 44
function/avro.json After 0 0 0 4 4
function/bitwise.json After 15 0 0 0 15
Before 15 0 0 0 15
function/collection.json After 12 0 0 1 13
Before 12 0 0 0 12
function/conditional.json After 15 0 0 0 15
Before 15 0 0 0 15
function/conversion.json After 2 0 0 0 2
Before 2 0 0 0 2
function/csv.json After 5 0 0 0 5
Before 5 0 0 0 5
function/datetime.json After 167 0 0 40 207
Before 167 0 0 13 180
function/generator.json After 13 0 0 0 13
Before 13 0 0 0 13
function/hash.json After 6 0 0 1 7
Before 6 0 0 1 7
function/json.json After 23 0 0 0 23
Before 22 0 0 0 22
function/lambda.json After 21 0 0 10 31
Before 21 0 0 10 31
function/map.json After 11 0 0 0 11
Before 11 0 0 0 11
function/math.json After 124 0 0 0 124
Before 124 0 0 0 124
function/misc.json After 33 0 0 13 46
Before 39 0 0 35 74
function/predicate.json After 72 0 0 7 79
Before 72 0 0 7 79
function/protobuf.json After 0 0 0 2 2
function/sketch.json After 6 0 0 35 41
function/st.json After 2 0 0 8 10
Before 2 0 0 5 7
function/string.json After 193 0 0 12 205
Before 193 0 0 12 205
function/struct.json After 2 0 0 0 2
Before 2 0 0 0 2
function/url.json After 10 0 0 0 10
Before 10 0 0 0 10
function/variant.json After 28 0 0 2 30
Before 28 0 0 0 28
function/vector.json After 0 0 0 11 11
function/window.json After 6 0 0 3 9
Before 6 0 0 3 9
function/xml.json After 15 0 0 2 17
Before 15 0 0 2 17
plan/ddl_alter_table.json After 49 14 3 11 77
Before 49 14 3 11 77
plan/ddl_alter_view.json After 5 1 0 0 6
Before 5 1 0 0 6
plan/ddl_analyze_table.json After 17 6 0 0 23
Before 17 6 0 0 23
plan/ddl_cache.json After 4 0 1 0 5
Before 4 0 1 0 5
plan/ddl_create_index.json After 0 0 0 3 3
Before 0 0 0 3 3
plan/ddl_create_table.json After 60 27 11 7 105
Before 60 27 11 7 105
plan/ddl_delete_from.json After 2 1 0 0 3
Before 2 1 0 0 3
plan/ddl_describe.json After 4 0 0 0 4
Before 7 0 0 0 7
plan/ddl_drop_index.json After 0 0 0 2 2
Before 0 0 0 2 2
plan/ddl_drop_view.json After 5 0 0 0 5
Before 5 0 0 0 5
plan/ddl_insert_into.json After 17 3 1 11 32
Before 16 1 1 0 18
plan/ddl_insert_overwrite.json After 9 0 2 0 11
Before 9 0 2 0 11
plan/ddl_load_data.json After 4 0 0 0 4
Before 4 0 0 0 4
plan/ddl_merge_into.json After 8 5 3 0 16
Before 8 4 3 0 15
plan/ddl_misc.json After 10 0 0 10 20
Before 10 0 0 0 10
plan/ddl_replace_table.json After 56 13 8 7 84
Before 56 13 8 7 84
plan/ddl_select.json After 1 0 0 0 1
Before 1 0 0 0 1
plan/ddl_show_views.json After 7 0 0 0 7
Before 7 0 0 0 7
plan/ddl_uncache.json After 2 0 0 0 2
Before 2 0 0 0 2
plan/ddl_update.json After 2 1 0 0 3
Before 2 1 0 0 3
plan/error_alter_table.json After 0 2 0 0 2
Before 0 2 0 0 2
plan/error_analyze_table.json After 0 1 0 0 1
Before 0 1 0 0 1
plan/error_create_table.json After 0 6 0 0 6
Before 0 6 0 0 6
plan/error_describe.json After 0 1 0 0 1
Before 0 1 0 0 1
plan/error_join.json After 0 2 0 0 2
Before 0 2 0 0 2
plan/error_load_data.json After 0 1 0 0 1
Before 0 1 0 0 1
plan/error_misc.json After 0 14 0 0 14
Before 0 14 0 0 14
plan/error_order_by.json After 1 4 0 0 5
Before 1 4 0 0 5
plan/error_select.json After 0 15 0 0 15
Before 0 15 0 0 15
plan/error_with.json After 0 1 0 0 1
Before 0 1 0 0 1
plan/plan_alter_view.json After 0 2 0 0 2
Before 0 2 0 0 2
plan/plan_create_view.json After 0 2 0 0 2
Before 0 2 0 0 2
plan/plan_explain.json After 0 1 1 0 2
Before 0 1 1 0 2
plan/plan_group_by.json After 9 1 0 1 11
Before 9 1 0 1 11
plan/plan_hint.json After 25 0 3 0 28
Before 25 0 3 0 28
plan/plan_insert_into.json After 3 0 0 0 3
Before 3 0 0 0 3
plan/plan_insert_overwrite.json After 2 0 0 0 2
Before 2 0 0 0 2
plan/plan_join.json After 59 9 1 4 73
Before 59 2 1 0 62
plan/plan_misc.json After 15 4 0 10 29
Before 15 4 0 10 29
plan/plan_order_by.json After 15 5 1 10 31
Before 15 5 1 10 31
plan/plan_select.json After 107 26 3 51 187
Before 85 14 5 16 120
plan/plan_set_operation.json After 17 0 0 0 17
Before 17 0 0 0 17
plan/plan_with.json After 6 0 2 0 8
Before 6 0 1 0 7
plan/unpivot_join.json After 4 0 0 0 4
Before 4 0 0 0 4
plan/unpivot_select.json After 14 6 0 0 20
Before 14 6 0 0 20
table_schema.json After 8 6 0 0 14
Before 8 6 0 0 14

@codecov

codecov Bot commented Jul 25, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.60773% with 17 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
crates/sail-python-udf/src/cereal/pyspark_udtf.rs 90.16% 6 Missing ⚠️
crates/sail-spark-connect/src/proto/plan.rs 80.00% 5 Missing ⚠️
crates/sail-python-udf/src/cereal/mod.rs 87.50% 3 Missing ⚠️
crates/sail-spark-connect/src/server.rs 0.00% 2 Missing ⚠️
crates/sail-execution/src/proto/codec.rs 66.66% 1 Missing ⚠️
@@            Coverage Diff             @@
##             main    #2289      +/-   ##
==========================================
- Coverage   79.46%   78.32%   -1.15%     
==========================================
  Files         909      916       +7     
  Lines      171774   174768    +2994     
==========================================
+ Hits       136503   136883     +380     
- Misses      35271    37885    +2614     
Flag Coverage Δ
catalog-integration-tests 25.36% <8.28%> (-3.25%) ⬇️
ibis-tests 16.92% <29.83%> (+0.01%) ⬆️
python-unit-tests 62.22% <84.53%> (+0.20%) ⬆️
rust-slow-tests 45.73% <6.62%> (-1.27%) ⬇️
rust-unit-tests 42.00% <6.62%> (-1.20%) ⬇️
spark-tests 31.93% <79.00%> (+0.20%) ⬆️
Files with missing lines Coverage Δ
crates/sail-plan/src/resolver/query/udtf.rs 94.61% <100.00%> (+0.03%) ⬆️
crates/sail-python-udf/src/cereal/pyspark_udf.rs 100.00% <100.00%> (ø)
crates/sail-python-udf/src/config.rs 90.62% <100.00%> (+0.96%) ⬆️
crates/sail-spark-connect/src/config.rs 90.46% <100.00%> (+0.56%) ⬆️
crates/sail-execution/src/proto/codec.rs 48.44% <66.66%> (+1.07%) ⬆️
crates/sail-spark-connect/src/server.rs 81.01% <0.00%> (-2.11%) ⬇️
crates/sail-python-udf/src/cereal/mod.rs 92.30% <87.50%> (+1.13%) ⬆️
crates/sail-spark-connect/src/proto/plan.rs 88.75% <80.00%> (+0.76%) ⬆️
crates/sail-python-udf/src/cereal/pyspark_udtf.rs 96.42% <90.16%> (-3.58%) ⬇️

... and 71 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

github-actions Bot commented Jul 25, 2026

Copy link
Copy Markdown

Spark 3.5.7 Test Report

Commit Information

Commit Revision Branch
After 57d3fb7 refs/pull/2289/merge
Before 6d1e068 refs/heads/main

Test Summary

Suite Commit Failed Passed Skipped Warnings Time (s)
doctest-catalog After 10 14 1 6 6.08
Before 10 14 1 6 5.90
doctest-column After 33 2 6.12
Before 33 2 5.45
doctest-dataframe After 14 83 10 3 8.64
Before 14 83 10 3 7.88
doctest-functions After 13 387 9 8 17.34
Before 13 387 9 7 15.54
test-connect After 111 890 170 707 145.15
Before 111 890 170 693 127.79

Test Details

Error Counts
          148 Total
           85 Total Unique
-------- ---- ----------------------------------------------------------------------------------------------------------
           13 PySparkAssertionError: [DIFFERENT_PANDAS_DATAFRAME] DataFrames are not almost equal:
           10 handle add artifacts
            9 DocTestFailure
            6 UnsupportedOperationException: PlanNode::CacheTable
            5 UnsupportedOperationException: function: input_file_name
            4 AssertionError: False is not true
            3 ValueError: Converting to Python dictionary is not supported when duplicate field names are present
            2 AnalysisException: Could not find config namespace "spark"
            2 AnalysisException: Internal error: Function 'approx_percentile_cont' failed to match any signature, errors: Error during planning: Function 'approx_percentile_cont' expects 2 arguments but received 3,...
            2 AnalysisException: No table format found for: orc
            2 AnalysisException: explode should be rewritten during logical plan analysis
            2 AnalysisException: not supported: function exists
            2 AssertionError
            2 AssertionError: 0 not greater than or equal to 1
            2 IllegalArgumentException: expected value at line 1 column 1
            2 IllegalArgumentException: invalid argument: found FUNCTION at 7:15 expected 'DATABASE', 'SCHEMA', 'NAMESPACE', 'OR', 'TEMP', 'TEMPORARY', 'EXTERNAL', 'TABLE', 'GLOBAL', or 'VIEW'
            2 IllegalArgumentException: invalid argument: found RESET at 0:5 expected something else, ';', statement, or end of input
            2 PySparkNotImplementedError: [NOT_IMPLEMENTED] rdd() is not implemented.
(+2)        2 UnsupportedOperationException: Aggregate can not be used as a sliding accumulator because `retract_batch` is not implemented: avg@5zn41om0e223ea0060za4fgfl(#9) PARTITION BY [#8] ORDER BY [#9 ASC NULLS...
            2 UnsupportedOperationException: approx quantile
            2 UnsupportedOperationException: collect metrics
            2 UnsupportedOperationException: freq items
            2 UnsupportedOperationException: function: session_window
            2 UnsupportedOperationException: handle analyze same semantics
            2 UnsupportedOperationException: user defined data type should only exist in a field
            2 UnsupportedOperationException: with watermark
            2 handle artifact statuses
            1 AnalysisException: Table already exists: tbl1
            1 AnalysisException: Temporary View not found: tab2
(+1)        1 AnalysisException: UNION queries have different number of columns: left has 2 columns whereas right has 3 columns
(+1)        1 AnalysisException: expected datetime literal '-' at byte offset 2
            1 AnalysisException: not supported: qualified function name
(+1)        1 AssertionError: "Database 'memory:0ca4d160-f32c-4fc5-9f40-474360fc10ab' dropped." does not match "No table format found for: jdbc. The JDBC data source is provided by pysail and must be registered bef...
(+1)        1 AssertionError: "Database 'memory:9dc9c38c-a60a-4174-be07-50ffbfec3c4c' dropped." does not match "No table format found for: jdbc. The JDBC data source is provided by pysail and must be registered bef...
            1 AssertionError: 1 != 0
            1 AssertionError: 4 != 10
            1 AssertionError: AnalysisException not raised
            1 AssertionError: AnalysisException not raised by <lambda>
            1 AssertionError: Exception not raised
            1 AssertionError: Exception not raised by <lambda>
            1 AssertionError: Lists differ: [Row([178 chars]on='<<'), Row(function='<='), Row(function='<=[13267 chars]'~')] != [Row([178 chars]on='<='), Row(function='<=>'), Row(function='<[11284 chars]'~')]
            1 AssertionError: Lists differ: [Row(id=90, name='90'), Row(id=91, name='91'), Ro[176 chars]99')] != [Row(id=15, name='15'), Row(id=16, name='16'), Ro[176 chars]24')]
            1 AssertionError: Lists differ: [Row(ln(id)=0.0, ln(id)=0.0, struct(id, name)=Row(id=[1232 chars]0'))] != [Row(ln(id)=4.31748811353631, ln(id)=4.31748811353631[1312 chars]4'))]
            1 AssertionError: Lists differ: [Row(name='Andy', age=30), Row(name='Andy', ag[374 chars]one)] != [Row(age=19, name='Justin'), Row(age=19, name=[374 chars]el')]
            1 AssertionError: Lists differ: [Row(name='Andy', age=30), Row(name='Justin', [34 chars]one)] != [Row(_corrupt_record=' "age":19}\n', name=None[104 chars]el')]
            1 AssertionError: Row(point='[1.0, 2.0]', pypoint='[3.0, 4.0]') != Row(point='(1.0, 2.0)', pypoint='[3.0, 4.0]')
            1 AssertionError: StorageLevel(False, True, True, False, 1) != StorageLevel(False, False, False, False, 1)
            1 AssertionError: Struc[15 chars]eld('a', NullType(), True), StructField('b', L[51 chars]ue)]) != Struc[15 chars]eld('b', LongType(), True), StructField('c', S[15 chars]ue)])
            1 AssertionError: Struc[30 chars]estampType(), True), StructField('val', IntegerType(), True)]) != Struc[30 chars]estampType(), True), StructField('val', IntegerType(), False)])
            1 AssertionError: Struc[32 chars]e(), False), StructField('b', DoubleType(), Fa[158 chars]ue)]) != Struc[32 chars]e(), True), StructField('b', DoubleType(), Tru[154 chars]ue)])
            1 AssertionError: Struc[40 chars]ue), StructField('val', ArrayType(DoubleType(), False), True)]) != Struc[40 chars]ue), StructField('val', PythonOnlyUDT(), True)])
            1 AssertionError: YearMonthIntervalType(0, 1) != YearMonthIntervalType(0, 0)
            1 AssertionError: [1.0, 2.0] != ExamplePoint(1.0,2.0)
            1 AssertionError: dtype('<M8[us]') != 'datetime64[ns]'
            1 IllegalArgumentException: invalid argument: field not found in input schema: col1
            1 IllegalArgumentException: invalid argument: table does not exist: ObjectName([Identifier("test_table")])
            1 PySparkNotImplementedError: [NOT_IMPLEMENTED] toJSON() is not implemented.
            1 PythonException:  AttributeError: 'NoneType' object has no attribute 'partitionId'
            1 PythonException:  AttributeError: 'list' object has no attribute 'y'
            1 SparkRuntimeException: Cast error: Cannot cast string '1997/02/28 10:30:00' to value of Date32 type
            1 SparkRuntimeException: Invalid argument error: 83.140 is too large to store in a Decimal128 of precision 4. Max is 9.999
            1 SparkRuntimeException: Json error: Not valid JSON: EOF while parsing a list at line 1 column 1
            1 SparkRuntimeException: Json error: Not valid JSON: expected value at line 1 column 2
            1 SparkRuntimeException: Parser error: Error while parsing value '0
            1 SparkRuntimeException: This feature is not implemented: Unsupported CAST from Map("entries": non-null Struct("key": non-null Int32, "value": non-null Int32), unsorted) to Null
(+1)        1 UnsupportedOperationException: Aggregate can not be used as a sliding accumulator because `retract_batch` is not implemented: avg@5zn41om0e223ea0060za4fgfl(#9) PARTITION BY [#8] ORDER BY [#9 ASC NULLS...
(+1)        1 UnsupportedOperationException: Aggregate can not be used as a sliding accumulator because `retract_batch` is not implemented: avg@5zn41om0e223ea0060za4fgfl(plus_one@rxpc2jkqe1qo369e7zvguy24(#9)) PARTI...
            1 UnsupportedOperationException: PlanNode::ClearCache
            1 UnsupportedOperationException: PlanNode::IsCached
            1 UnsupportedOperationException: PlanNode::RecoverPartitions
            1 UnsupportedOperationException: Support for 'approx_distinct' for data type Float64 is not implemented
            1 UnsupportedOperationException: apply in pandas with state
            1 UnsupportedOperationException: bucketing for writing listing table format
            1 UnsupportedOperationException: deduplicate within watermark
            1 UnsupportedOperationException: function: input_file_block_length
            1 UnsupportedOperationException: function: input_file_block_start
            1 UnsupportedOperationException: function: map_filter
            1 UnsupportedOperationException: function: map_zip_with
            1 UnsupportedOperationException: function: transform_keys
            1 UnsupportedOperationException: function: transform_values
            1 UnsupportedOperationException: function: zip_with
            1 UnsupportedOperationException: handle analyze semantic hash
            1 UnsupportedOperationException: unknown function: distributed_sequence_id
            1 ValueError: The column label 'id' is not unique.
            1 ValueError: The column label 'struct' is not unique.
(-1)        0 AnalysisException: UNION queries have different number of columns: left has 3 columns whereas right has 2 columns
(-1)        0 AssertionError: "Database 'memory:1552b7d4-d8d7-416c-9853-0b4ebf3e9e9a' dropped." does not match "No table format found for: jdbc. The JDBC data source is provided by pysail and must be registered bef...
(-1)        0 AssertionError: "Database 'memory:254ef398-d0e6-47d9-b704-43abc370dca7' dropped." does not match "No table format found for: jdbc. The JDBC data source is provided by pysail and must be registered bef...
(-1)        0 PythonException:  AttributeError: 'list' object has no attribute 'x'
(-2)        0 UnsupportedOperationException: Aggregate can not be used as a sliding accumulator because `retract_batch` is not implemented: avg@bjwplt75b8s8lnn9byry0pm4s(#9) PARTITION BY [#8] ORDER BY [#9 ASC NULLS...
(-1)        0 UnsupportedOperationException: Aggregate can not be used as a sliding accumulator because `retract_batch` is not implemented: avg@bjwplt75b8s8lnn9byry0pm4s(#9) PARTITION BY [#8] ORDER BY [#9 ASC NULLS...
(-1)        0 UnsupportedOperationException: Aggregate can not be used as a sliding accumulator because `retract_batch` is not implemented: avg@bjwplt75b8s8lnn9byry0pm4s(plus_one@5u7wzy1x1zf4loxm9n1r277lw(#9)) PART...
Passed Tests Diff
--- before.txt	2026-07-27 03:49:43.685326834 +0000
+++ after.txt	2026-07-27 03:49:43.873329920 +0000
@@ -925 +924,0 @@
-pyspark/sql/tests/connect/test_parity_errors.py::ErrorsParityTests::test_date_time_exception
@@ -1152,0 +1152 @@
+pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_complex_nested_udt_in_df
Failed Tests
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.cacheTable
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.clearCache
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.createTable
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.functionExists
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.getFunction
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.isCached
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.recoverPartitions
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.refreshByPath
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.refreshTable
pyspark/sql/catalog.py::pyspark.sql.catalog.Catalog.uncacheTable
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.colRegex
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.dropDuplicatesWithinWatermark
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.explain
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.hint
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.observe
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.randomSplit
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.repartition
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.repartitionByRange
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.sameSemantics
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.sampleBy
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.storageLevel
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.toJSON
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrame.withWatermark
pyspark/sql/dataframe.py::pyspark.sql.dataframe.DataFrameStatFunctions.sampleBy
pyspark/sql/functions.py::pyspark.sql.functions.approx_percentile
pyspark/sql/functions.py::pyspark.sql.functions.input_file_block_length
pyspark/sql/functions.py::pyspark.sql.functions.input_file_block_start
pyspark/sql/functions.py::pyspark.sql.functions.input_file_name
pyspark/sql/functions.py::pyspark.sql.functions.map_entries
pyspark/sql/functions.py::pyspark.sql.functions.map_filter
pyspark/sql/functions.py::pyspark.sql.functions.map_zip_with
pyspark/sql/functions.py::pyspark.sql.functions.percentile_approx
pyspark/sql/functions.py::pyspark.sql.functions.regexp_instr
pyspark/sql/functions.py::pyspark.sql.functions.session_window
pyspark/sql/functions.py::pyspark.sql.functions.transform_keys
pyspark/sql/functions.py::pyspark.sql.functions.transform_values
pyspark/sql/functions.py::pyspark.sql.functions.zip_with
pyspark/sql/tests/connect/client/test_artifact.py::ArtifactTests::test_add_archive
pyspark/sql/tests/connect/client/test_artifact.py::ArtifactTests::test_add_file
pyspark/sql/tests/connect/client/test_artifact.py::ArtifactTests::test_add_pyfile
pyspark/sql/tests/connect/client/test_artifact.py::ArtifactTests::test_add_zipped_package
pyspark/sql/tests/connect/client/test_artifact.py::ArtifactTests::test_basic_requests
pyspark/sql/tests/connect/client/test_artifact.py::ArtifactTests::test_cache_artifact
pyspark/sql/tests/connect/client/test_artifact.py::ArtifactTests::test_copy_from_local_to_fs
pyspark/sql/tests/connect/client/test_artifact.py::LocalClusterArtifactTests::test_add_archive
pyspark/sql/tests/connect/client/test_artifact.py::LocalClusterArtifactTests::test_add_file
pyspark/sql/tests/connect/client/test_artifact.py::LocalClusterArtifactTests::test_add_pyfile
pyspark/sql/tests/connect/client/test_artifact.py::LocalClusterArtifactTests::test_add_zipped_package
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_collect
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_collect_timestamp
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_column_regexp
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_create_global_temp_view
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_deduplicate_within_watermark_in_batch
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_describe
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_hint
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_join_hint
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_json
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_multi_paths
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_observe
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_orc
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_random_split
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_same_semantics
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_schema
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_semantic_hash
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_simple_udt_from_read
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_sql_with_command
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_stat_approx_quantile
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_stat_freq_items
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_stat_sample_by
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_streaming_local_relation
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_tail
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_to
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_with_local_list
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_with_local_ndarray
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectBasicTests::test_write_operations
pyspark/sql/tests/connect/test_connect_basic.py::SparkConnectSessionTests::test_error_stack_trace
pyspark/sql/tests/connect/test_connect_column.py::SparkConnectColumnTests::test_column_arithmetic_ops
pyspark/sql/tests/connect/test_connect_column.py::SparkConnectColumnTests::test_decimal
pyspark/sql/tests/connect/test_connect_column.py::SparkConnectColumnTests::test_distributed_sequence_id
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_aggregation_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_collection_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_date_ts_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_generator_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_lambda_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_map_collection_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_math_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_normal_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_string_functions_multi_args
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_time_window_functions
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_udf
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_udtf
pyspark/sql/tests/connect/test_connect_function.py::SparkConnectFunctionTests::test_window_functions
pyspark/sql/tests/connect/test_parity_arrow.py::ArrowParityTests::test_createDataFrame_duplicate_field_names
pyspark/sql/tests/connect/test_parity_arrow.py::ArrowParityTests::test_pandas_self_destruct
pyspark/sql/tests/connect/test_parity_arrow.py::ArrowParityTests::test_toPandas_duplicate_field_names
pyspark/sql/tests/connect/test_parity_arrow_python_udf.py::ArrowPythonUDFParityTests::test_udf_with_input_file_name
pyspark/sql/tests/connect/test_parity_arrow_python_udf.py::UDFParityTests::test_udf_with_input_file_name
pyspark/sql/tests/connect/test_parity_catalog.py::CatalogParityTests::test_function_exists
pyspark/sql/tests/connect/test_parity_catalog.py::CatalogParityTests::test_get_function
pyspark/sql/tests/connect/test_parity_catalog.py::CatalogParityTests::test_list_functions
pyspark/sql/tests/connect/test_parity_catalog.py::CatalogParityTests::test_refresh_table
pyspark/sql/tests/connect/test_parity_catalog.py::CatalogParityTests::test_table_cache
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_cache_dataframe
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_cache_table
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_duplicate_field_names
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_extended_hint_types
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_freqItems
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_generic_hints
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_input_files
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_to
pyspark/sql/tests/connect/test_parity_dataframe.py::DataFrameParityTests::test_to_pandas
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_checking_csv_header
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_encoding_json
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_ignore_column_of_all_nulls
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_jdbc
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_jdbc_format
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_linesep_json
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_multiline_json
pyspark/sql/tests/connect/test_parity_datasources.py::DataSourcesParityTests::test_read_multiple_orc_file
pyspark/sql/tests/connect/test_parity_errors.py::ErrorsParityTests::test_date_time_exception
pyspark/sql/tests/connect/test_parity_functions.py::FunctionsParityTests::test_approxQuantile
pyspark/sql/tests/connect/test_parity_functions.py::FunctionsParityTests::test_functions_broadcast
pyspark/sql/tests/connect/test_parity_functions.py::FunctionsParityTests::test_input_file_name_udf
pyspark/sql/tests/connect/test_parity_pandas_grouped_map.py::GroupedApplyInPandasTests::test_grouped_over_window
pyspark/sql/tests/connect/test_parity_pandas_grouped_map.py::GroupedApplyInPandasTests::test_grouped_over_window_with_key
pyspark/sql/tests/connect/test_parity_pandas_grouped_map_with_state.py::GroupedApplyInPandasWithStateTests::test_apply_in_pandas_with_state_python_worker_random_failure
pyspark/sql/tests/connect/test_parity_pandas_udf_scalar.py::PandasUDFScalarParityTests::test_scalar_iter_udf_init
pyspark/sql/tests/connect/test_parity_pandas_udf_scalar.py::PandasUDFScalarParityTests::test_vectorized_udf_check_config
pyspark/sql/tests/connect/test_parity_pandas_udf_scalar.py::PandasUDFScalarParityTests::test_vectorized_udf_invalid_length
pyspark/sql/tests/connect/test_parity_pandas_udf_window.py::PandasUDFWindowParityTests::test_bounded_mixed
pyspark/sql/tests/connect/test_parity_pandas_udf_window.py::PandasUDFWindowParityTests::test_bounded_simple
pyspark/sql/tests/connect/test_parity_pandas_udf_window.py::PandasUDFWindowParityTests::test_shrinking_window
pyspark/sql/tests/connect/test_parity_pandas_udf_window.py::PandasUDFWindowParityTests::test_sliding_window
pyspark/sql/tests/connect/test_parity_readwriter.py::ReadwriterParityTests::test_bucketed_write
pyspark/sql/tests/connect/test_parity_readwriter.py::ReadwriterParityTests::test_save_and_load
pyspark/sql/tests/connect/test_parity_readwriter.py::ReadwriterParityTests::test_save_and_load_builder
pyspark/sql/tests/connect/test_parity_readwriter.py::ReadwriterV2ParityTests::test_table_overwrite
pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_cast_to_string_with_udt
pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_cast_to_udt_with_udt
pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_negative_decimal
pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_parquet_with_udt
pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_udf_with_udt
pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_udt_with_none
pyspark/sql/tests/connect/test_parity_types.py::TypesParityTests::test_yearmonth_interval_type
pyspark/sql/tests/connect/test_parity_udf.py::UDFParityTests::test_udf_with_input_file_name
pyspark/sql/tests/connect/test_parity_udtf.py::ArrowUDTFParityTests::test_udtf_arrow_sql_conf
pyspark/sql/tests/connect/test_utils.py::ConnectUtilsTests::test_assert_approx_equal_decimaltype_custom_rtol_pass
pyspark/sql/tests/connect/test_utils.py::ConnectUtilsTests::test_assert_equal_nested_struct_str_duplicate

@github-actions

github-actions Bot commented Jul 25, 2026

Copy link
Copy Markdown

Ibis Test Report

Commit Information

Commit Revision Branch
After 57d3fb7 refs/pull/2289/merge
Before 6d1e068 refs/heads/main

Test Summary

Suite Commit Failed Passed Skipped Warnings Time (s)
test-ibis After 32 1537 166 4535 238.16
Before 34 1535 166 4524 237.37

Test Details

Error Counts
(-2)       33 Total
(-2)       24 Total Unique
-------- ---- ----------------------------------------------------------------------------------------------------------
            5 IllegalArgumentException: invalid argument: found TRUNCATE at 0:8 expected something else, ';', statement, or end of input
            2 AnalysisException: Internal error: Function 'approx_percentile_cont' failed to match any signature, errors: Error during planning: Function 'approx_percentile_cont' requires Float64, but received List...
            2 AssertionError
            2 AssertionError: Series are different
            2 IllegalArgumentException: invalid argument: found PARTITIONS at 5:15 expected 'DATABASES', 'SCHEMAS', 'NAMESPACES', 'CATALOGS', 'TABLES', 'TABLE', 'CREATE', 'COLUMNS', 'VIEWS', 'ALL', 'USER', 'SYSTEM'...
            2 assert ibis.Schema {... float64\n} == ibis.Schema {... float64\n} Full diff: ibis.Schema { carat float64 cut string color string clarity string depth float64 table float64 - price int32 ? ^^ + price i...
            1 AnalysisException: Catalog not found: local
(+1)        1 AnalysisException: Database not found: ibis_database_el5la3l3azfvxldpbnfdbjluka
            1 AssertionError: DataFrame.iloc[:, 0] (column name="id") are different
            1 AssertionError: DataFrame.iloc[:, 0] (column name="playerID") are different
            1 AssertionError: Series NA mask are different
(+1)        1 AssertionError: assert 'ibis_cached_hw23cbbrffhy7job2shyld7jfe' not in ['array_types', 'astronauts', 'awards_players', 'basic_table', 'batting', 'complicated', ...]
            1 AssertionError: assert nan == 22
            1 Failed: DID NOT RAISE <class 'pyspark.errors.exceptions.base.AnalysisException'>
            1 SparkRuntimeException: Cast error: Casting from Date32 to Float64 not supported
            1 SparkRuntimeException: Error during planning: expr type Struct("StructColumn({'x': xs, 'y': ys})": non-null Struct("x": non-null Int32, "y": non-null Int32)) can't cast to Struct("x": Int64, metadata:...
            1 TypeError: Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array
            1 UnsupportedOperationException: CommandNode::AnalyzeTable
            1 UnsupportedOperationException: Physical plan does not support logical expression AggregateFunction(AggregateFunction { func: AggregateUDF { inner: ArrayAgg { signature: Signature { type_signature: Any...
            1 UnsupportedOperationException: Physical plan does not support logical expression InSubquery(InSubquery { expr: Column(Column { relation: Some(Bare { table: "t0" }), name: "#0" }), subquery: <subquery>...
            1 UnsupportedOperationException: Physical plan does not support logical expression InSubquery(InSubquery { expr: Column(Column { relation: Some(Bare { table: "t0" }), name: "#1" }), subquery: <subquery>...
            1 UnsupportedOperationException: unsupported ALTER TABLE operation
            1 assert frozenset({None}) == frozenset({None, 47}) Extra items in the right set: 47 Full diff: frozenset({ None, - 47, })
            1 assert {0.0, 1.0, 2.0, 3.0} == {1, 2, 3} Extra items in the left set: 0.0 Full diff: { + 0.0, - 1, + 1.0, ? ++ - 2, + 2.0, ? ++ - 3, + 3.0, ? ++ }
(-1)        0 AnalysisException: Database not found: ibis_database_d645dgukp5hwtbbu3an35wbrd4
(-1)        0 AssertionError: assert 'ibis_cached_lwtflfywbzhn5pibox7g2vrntm' not in ['array_types', 'astronauts', 'awards_players', 'basic_table', 'batting', 'complicated', ...]
(-1)        0 DateTimeException: Error parsing timestamp from '01/01/09' using format '%-%-M/%-d/%y': bad or unsupported format string
(-1)        0 ValueError: NaTType does not support strftime
Passed Tests Diff
--- before.txt	2026-07-27 03:50:55.168126097 +0000
+++ after.txt	2026-07-27 03:50:55.397245999 +0000
@@ -1431,0 +1432 @@
+ibis/backends/tests/test_temporal.py::test_string_as_date[pyspark-mysql_format]
@@ -1432,0 +1434 @@
+ibis/backends/tests/test_temporal.py::test_string_as_timestamp[pyspark-mysql_format]
Failed Tests
ibis/backends/pyspark/tests/test_basic.py::test_group_by
ibis/backends/pyspark/tests/test_client.py::test_catalog_db_args
ibis/backends/pyspark/tests/test_client.py::test_create_table_with_partition_and_catalog
ibis/backends/pyspark/tests/test_client.py::test_create_table_with_partition_no_catalog
ibis/backends/pyspark/tests/test_ddl.py::test_compute_stats
ibis/backends/pyspark/tests/test_ddl.py::test_drop_non_empty_database
ibis/backends/pyspark/tests/test_ddl.py::test_insert_table
ibis/backends/pyspark/tests/test_ddl.py::test_truncate_table
ibis/backends/tests/test_aggregation.py::test_approx_quantile[pyspark-True-False]
ibis/backends/tests/test_aggregation.py::test_approx_quantile[pyspark-True-True]
ibis/backends/tests/test_aggregation.py::test_date_quantile[pyspark]
ibis/backends/tests/test_aggregation.py::test_group_concat_over_window[pyspark]
ibis/backends/tests/test_client.py::test_insert_overwrite_from_dataframe[pyspark]
ibis/backends/tests/test_client.py::test_insert_overwrite_from_expr[pyspark]
ibis/backends/tests/test_client.py::test_insert_overwrite_from_list[pyspark]
ibis/backends/tests/test_client.py::test_rename_table[pyspark]
ibis/backends/tests/test_export.py::test_table_to_csv[pyspark]
ibis/backends/tests/test_expr_caching.py::test_persist_expression_contextmanager[pyspark]
ibis/backends/tests/test_expr_caching.py::test_persist_expression_release[pyspark]
ibis/backends/tests/test_expr_caching.py::test_persist_expression_repeated_cache[pyspark]
ibis/backends/tests/test_generic.py::test_isin_uncorrelated[pyspark]
ibis/backends/tests/test_generic.py::test_isin_uncorrelated_simple[pyspark]
ibis/backends/tests/test_io.py::test_read_csv[pyspark-default]
ibis/backends/tests/test_io.py::test_read_csv[pyspark-file_name]
ibis/backends/tests/test_join.py::test_join_with_pandas[pyspark]
ibis/backends/tests/test_json.py::test_json_getitem_array[pyspark]
ibis/backends/tests/test_struct.py::test_field_overwrite_always_prefers_unpacked[pyspark]
ibis/backends/tests/test_struct.py::test_isin_struct[pyspark]
ibis/backends/tests/test_struct.py::test_single_field[pyspark-a]
ibis/backends/tests/test_struct.py::test_single_field[pyspark-b]
ibis/backends/tests/test_struct.py::test_single_field[pyspark-c]
ibis/backends/tests/test_temporal.py::test_delta[pyspark-time]

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Updates Sail’s Spark Connect compatibility layer, test harnesses, and documentation to support Apache Spark / PySpark 4.2.0, including protocol (proto) additions and adjustments to Python UDF/UDTF serialization behavior needed for Spark 4.2 workers.

Changes:

  • Bump Spark/PySpark versions across CI, packaging, Docker images, and docs from 4.1.1 → 4.2.0 (and expand the test matrix accordingly).
  • Update Spark Connect protocol definitions and server stubs to compile against Spark 4.2 proto additions (e.g., GetStatus, new relation/command variants).
  • Refresh Spark “gold data” and PySpark-related tests to reflect Spark 4.2 behavior and error semantics.

Reviewed changes

Copilot reviewed 59 out of 59 changed files in this pull request and generated no comments.

Show a summary per file
File Description
scripts/spark-tests/spark-4.2.0.patch Updates the maintained Spark-side patch for building/packaging a patched PySpark 4.2.0 used in integration tests.
scripts/spark-gold-data/spark-4.2.0.patch Updates the Spark-side patch used to regenerate/collect gold data against Spark 4.2.0.
scripts/spark-gold-data/bootstrap.sh Switches gold-data bootstrap to apply Spark v4.2.0 patch.
python/pysail/tests/spark/udf/test_udf_kwargs.py Adjusts expected exception types/messages for Spark 4.2 UDTF error behavior.
python/pysail/tests/spark/datasource/test_python.py Updates datasource partition expectations for Spark 4.2 (InputPartition default partition semantics).
python/pysail/tests/spark/dataframe/udt.py Adds importable UDT definitions for doctests (Spark 4.2 UDT resolution behavior).
python/pysail/tests/spark/dataframe/test_udt.txt Updates doctest to import UDTs from a module rather than defining them inline.
python/pysail/tests/spark/catalog/hms/conftest.py Skips HMS Delta-related tests when Spark minor isn’t mapped (instead of hard error).
pyproject.toml Bumps PySpark dependencies to 4.2.0 and extends Hatch test matrices for Spark 4.2.0.
docs/introduction/getting-started/index.md Updates getting-started install snippets to Spark/PySpark 4.2.0.
docs/guide/integrations/mcp-server.md Updates MCP install command example to use pyspark-client 4.2.0.
docs/development/spark-tests/test-spark.md Updates Spark test instructions/examples to refer to Spark 4.2.0 envs.
docs/development/spark-tests/spark-setup.md Updates Spark clone/build examples to use Spark v4.2.0.
docs/development/spark-tests/spark-patch.md Updates patch apply/diff/revert examples to Spark 4.2.0 patch file.
docs/development/build/python.md Updates example Hatch env invocation to Spark 4.2.0.
docker/release/Dockerfile Updates default PYSPARK_VERSION build arg to 4.2.0.
docker/quickstart/Dockerfile Updates default PYSPARK_VERSION build arg to 4.2.0.
docker/dev/Dockerfile Updates default PYSPARK_VERSION build arg to 4.2.0.
crates/sail-spark-connect/tests/gold_data/plan/plan_with.json Updates/extends plan gold data for Spark 4.2 parsing/planning behavior.
crates/sail-spark-connect/tests/gold_data/plan/plan_join.json Updates join parsing/planning gold data (incl. new/changed Spark 4.2 syntax/errors).
crates/sail-spark-connect/tests/gold_data/plan/error_misc.json Updates error message gold data to Spark 4.2 error classes/messages.
crates/sail-spark-connect/tests/gold_data/plan/ddl_misc.json Adds/updates DDL gold cases reflecting Spark 4.2 syntax support changes.
crates/sail-spark-connect/tests/gold_data/plan/ddl_merge_into.json Adds merge error-case gold data aligned with Spark 4.2 behavior.
crates/sail-spark-connect/tests/gold_data/plan/ddl_insert_into.json Adds/updates insert/replace planning gold data for Spark 4.2.
crates/sail-spark-connect/tests/gold_data/plan/ddl_describe.json Updates describe gold data (removes cases not represented the same way in Spark 4.2).
crates/sail-spark-connect/tests/gold_data/function/vector.json Adds Spark 4.2 function gold coverage entries for vector functions (unsupported in Sail).
crates/sail-spark-connect/tests/gold_data/function/variant.json Extends variant-related function gold data for Spark 4.2.
crates/sail-spark-connect/tests/gold_data/function/st.json Updates spatial function gold data to Spark 4.2 signatures/printing.
crates/sail-spark-connect/tests/gold_data/function/sketch.json Adds sketch-related function gold coverage (many unsupported in Sail).
crates/sail-spark-connect/tests/gold_data/function/protobuf.json Adds protobuf function gold coverage (not implemented in Sail).
crates/sail-spark-connect/tests/gold_data/function/misc.json Reorganizes misc function gold data (splitting/moving many function tests).
crates/sail-spark-connect/tests/gold_data/function/json.json Adds to_json(..., sortKeys) gold case for Spark 4.2.
crates/sail-spark-connect/tests/gold_data/function/datetime.json Adds Spark 4.2 time_bucket and time_* conversion function gold coverage (unsupported).
crates/sail-spark-connect/tests/gold_data/function/collection.json Adds binary reverse(x'...') gold case (shows current execution limitation).
crates/sail-spark-connect/tests/gold_data/function/avro.json Adds avro function gold coverage (not implemented).
crates/sail-spark-connect/tests/gold_data/function/agg.json Adds/updates aggregate function gold cases for Spark 4.2.
crates/sail-spark-connect/tests/gold_data/data_type.json Adds TIMESTAMP WITH LOCAL TIME ZONE type parsing gold cases.
crates/sail-spark-connect/src/server.rs Adds Spark Connect GetStatus RPC stub required by Spark 4.2 service definition.
crates/sail-spark-connect/src/proto/plan.rs Handles new Spark 4.2 proto fields/relations/commands by mapping or returning “unsupported”.
crates/sail-spark-connect/src/config.rs Adds Spark 4.2-related config keys and version-sensitive defaults and parsing.
crates/sail-spark-connect/proto/spark/connect/relations.proto Updates Spark Connect relation protos for 4.2 (e.g., RelationChanges, NearestByJoin, parse XML, source_name).
crates/sail-spark-connect/proto/spark/connect/pipelines.proto Updates pipeline protos for 4.2 (new command variants, identifiers, deprecations).
crates/sail-spark-connect/proto/spark/connect/commands.proto Adds Spark 4.2 write fields and stream options (e.g., schema evolution flag).
crates/sail-spark-connect/proto/spark/connect/catalog.proto Adds Spark 4.2 catalog command messages.
crates/sail-spark-connect/proto/spark/connect/base.proto Adds Spark 4.2 GetStatus request/response and observed-metrics error fields.
crates/sail-python-udf/src/python/spark.py Updates Arrow/Pandas conversion paths and worker-call conventions for PySpark 4.2.
crates/sail-python-udf/src/config.rs Adds a new UDF config toggle for Spark 4.2 pandas int extension dtype preference.
crates/sail-python-udf/src/cereal/pyspark_udtf.rs Updates UDTF payload serialization to match Spark 4.2 worker protocol (RunnerConf/EvalConf, conf maps).
crates/sail-python-udf/src/cereal/pyspark_udf.rs Updates UDF payload serialization for Spark 4.2 worker protocol changes.
crates/sail-python-udf/src/cereal/mod.rs Extends supported PySpark version detection to include 4.2 and adds shared conf writer.
crates/sail-execution/src/proto/codec.rs Extends remote execution codec to carry the new PySpark UDF config field.
crates/sail-execution/proto/sail/plan/physical.proto Extends physical plan proto to include the new PySpark UDF config field.
.github/workflows/spark-package-artifacts.yml Updates Spark artifact packaging workflow matrix to Spark 4.2.0.
.github/workflows/report.yml Updates report artifact naming to Spark 4.2.0.
.github/workflows/gold-data-script-validation.yml Pins gold-data validation to Spark v4.2.0.
.github/workflows/catalog-tests.yml Runs catalog tests under Spark 4.2.0 Hatch env.
.github/workflows/build.yml Updates build/test matrices to include Spark 4.2.0.

@linhr
linhr marked this pull request as ready for review July 26, 2026 15:55
@linhr linhr added run spark tests Trigger Spark tests on a pull request run ibis tests Trigger Ibis tests on a pull request run catalog tests Trigger catalog tests on a pull request labels Jul 26, 2026
@linhr
linhr merged commit 5689106 into main Jul 27, 2026
23 checks passed
@linhr
linhr deleted the spark-4-2 branch July 27, 2026 04:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run catalog tests Trigger catalog tests on a pull request run ibis tests Trigger Ibis tests on a pull request run spark tests Trigger Spark tests on a pull request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants