adriangb commented on PR #11038:
URL: https://github.com/apache/arrow-rs/pull/11038#issuecomment-5872798297

   @Jefffrey thanks for the review. I applied both suggestions and added the 
note about `None` in 63df2e0329.
   
   Regarding Spark: indeed it does what Java does. I ran Spark 4.2.0 locally on 
the same cases as the table in the PR description. `CAST(TIMESTAMP_NTZ ... AS 
TIMESTAMP)`, `to_utc_timestamp`, `convert_timezone` and a string cast all give 
the same result, with ANSI mode on and off. All times are UTC:
   
   | Case | Spark 4.2.0 | PostgreSQL 17 / DuckDB / this PR |
   | --- | --- | --- |
   | `America/New_York` `2024-11-03 01:30` — **ambiguous** | `05:30` (earlier) 
| `06:30` (later) |
   | `America/Havana` `2024-11-03 00:00` — **ambiguous** | `04:00` (earlier) | 
`05:00` (later) |
   | `Australia/Sydney` `2024-04-07 02:30` — **ambiguous** | `2024-04-06 15:30` 
(earlier) | `2024-04-06 16:30` (later) |
   | `America/New_York` `2024-03-10 02:30` — gap | `07:30` | `07:30` |
   | `America/Sao_Paulo` `2018-11-04 00:00` — gap | `03:00` | `03:00` |
   | `Australia/Sydney` `2024-10-06 02:30` — gap | `2024-10-05 16:30` | 
`2024-10-05 16:30` |
   | `Australia/Lord_Howe` `2024-10-06 02:15` — gap | `2024-10-05 15:45` | 
`2024-10-05 15:45` |
   | `Pacific/Chatham` `2024-09-29 03:00` — gap | `2024-09-28 14:15` | 
`2024-09-28 14:15` |
   
   So Spark agrees on every gap and disagrees on every ambiguous reading. The 
source confirms it: the cast calls 
[`convertTz`](https://github.com/apache/spark/blob/b1b2685058d6359e00842cbd4c723f4ce3810709/sql/catalyst/src/main/scala/org/apache/spark/sql/catalyst/expressions/Cast.scala#L901),
 which calls 
[`LocalDateTime.atZone`](https://github.com/apache/spark/blob/b1b2685058d6359e00842cbd4c723f4ce3810709/sql/api/src/main/scala/org/apache/spark/sql/catalyst/util/SparkDateTimeUtils.scala#L198).
 That is `ZonedDateTime.of`, which takes the earlier offset in an overlap and 
shifts forward in a gap. Spark never raises an error or returns NULL for these 
readings.
   
   So the SQL engines split on the ambiguous case: PostgreSQL and DuckDB take 
the later instant, Spark takes the earlier one. But, as far as I was able to 
tell, no Spark-compatible DataFusion implementer reaches this code. Comet 
handles the naive-to-zoned cast in its own code 
([cast.rs](https://github.com/apache/datafusion-comet/blob/764936187/native/spark-expr/src/conversion_funcs/cast.rs#L395)),
 which takes the earlier instant. That code runs before Comet falls back to 
arrow's cast. In datafusion-spark, to_utc_timestamp, from_utc_timestamp and 
spark_cast do not call it either. And compared with Spark, readings in a gap go 
from an error or NULL to Spark's answer. Ambiguous readings did not give 
Spark's answer before this PR either. So this PR does not make anything worse 
for Spark users, because no Spark-compatible engine reaches this kernel for 
this cast. 


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to