Description
After upgrading spark-bigquery-connector from 0.43.1 to a newer version,
reading data back from BigQuery produces two type regressions.
Regression 1: MapType deserialization broken
When writing a Spark DataFrame with MapType(StringType(), StringType()),
the connector correctly converts it to ARRAY<STRUCT<key STRING, value STRING>>
in BigQuery (expected, since BQ has no native MAP type).
However, on read, version 0.43.1 used to reconstruct the Spark MapType
and return Python dict values. The new version returns
array<struct<key:string,value:string>> - a list of Row objects - instead.
Expected (0.43.1):
tags={'a': 'b', 'c': 'd'}
Actual (new version):
tags=[Row(key='a', value='b'), Row(key='c', value='d')]
Read schema shows tags: array<struct<key:string,value:string>> instead of
tags: map<string,string>.
Regression 2: Decimal precision lost on read
Writing DecimalType(12,2) creates a NUMERIC(12, 2) column in BigQuery.
On read, 0.43.1 returned decimal(12,2). The new version returns decimal(38,9),
losing the original precision and scale, and producing values like
Decimal('123.450000000') instead of Decimal('123.45').
Steps to reproduce
- Write a Spark DataFrame with a
MapType(StringType(), StringType()) column
and a DecimalType(12,2) column to BigQuery via the connector
- Read back using
spark.read.format("bigquery").load(...)
- Observe schema and values differ from what was written
Impact
Any code that writes MapType columns and reads them back will observe
silent data shape change. Decimal values with explicit precision will
have incorrect scale on read.
Description
After upgrading spark-bigquery-connector from 0.43.1 to a newer version,
reading data back from BigQuery produces two type regressions.
Regression 1: MapType deserialization broken
When writing a Spark DataFrame with
MapType(StringType(), StringType()),the connector correctly converts it to
ARRAY<STRUCT<key STRING, value STRING>>in BigQuery (expected, since BQ has no native MAP type).
However, on read, version 0.43.1 used to reconstruct the Spark
MapTypeand return Python
dictvalues. The new version returnsarray<struct<key:string,value:string>>- a list ofRowobjects - instead.Expected (0.43.1):
tags={'a': 'b', 'c': 'd'}
Actual (new version):
tags=[Row(key='a', value='b'), Row(key='c', value='d')]
Read schema shows
tags: array<struct<key:string,value:string>>instead oftags: map<string,string>.Regression 2: Decimal precision lost on read
Writing
DecimalType(12,2)creates aNUMERIC(12, 2)column in BigQuery.On read, 0.43.1 returned
decimal(12,2). The new version returnsdecimal(38,9),losing the original precision and scale, and producing values like
Decimal('123.450000000')instead ofDecimal('123.45').Steps to reproduce
MapType(StringType(), StringType())columnand a
DecimalType(12,2)column to BigQuery via the connectorspark.read.format("bigquery").load(...)Impact
Any code that writes
MapTypecolumns and reads them back will observesilent data shape change. Decimal values with explicit precision will
have incorrect scale on read.