viirya commented on code in PR #56334:
URL: https://github.com/apache/spark/pull/56334#discussion_r3554432540
##########
sql/api/src/main/scala/org/apache/spark/sql/util/ArrowUtils.scala:
##########
@@ -38,6 +38,50 @@ private[sql] object ArrowUtils {
// todo: support more types.
+ /**
+ * Check if a Spark DataType is supported by Arrow. This recursively checks
complex types
+ * (Array, Struct, Map).
+ *
+ * Note: This checks compatibility with toArrowField(), not toArrowType().
Types like
+ * GeometryType, GeographyType, and VariantType are not supported by
toArrowType() (which only
+ * handles primitive Arrow types), but ARE supported by toArrowField() which
converts them to
+ * Arrow Struct representations with metadata. Since Arrow cache uses
toArrowField() via
+ * toArrowSchema() to create the schema, these types are supported.
+ */
+ def isSupportedByArrow(dt: DataType): Boolean = {
+ dt match {
+ // Primitive types
+ case BooleanType | ByteType | ShortType | IntegerType | LongType |
FloatType | DoubleType |
+ _: StringType | BinaryType | NullType =>
+ true
+
+ // Decimal
+ case _: DecimalType => true
+
+ // Temporal types
+ case DateType | TimestampType | TimestampNTZType | _: TimeType => true
Review Comment:
Done. The cache now stores nanosecond timestamps in SPARK-57975's lossless
struct encoding, so the full 0001-9999 domain round-trips -- including the
`9999-12-31T23:59:59.999999999` value from your repro -- restoring value-domain
parity with the default serializer. Wiring details: the stats collector reads
the struct components directly and compares with `TimestampNanosVal` ordering
(which also removes the epoch-nanos conversion, and with it the
precision-truncation asymmetry MaxGekk flagged); the zero-copy input
eligibility check now routes interchange-shaped int64-nanos vectors (e.g. a
Python UDF output feeding a cache) through the row-conversion path, like the
existing large-var-types handling; both reader modes are covered by the
existing type round-trip machinery. The docs' value-range note is updated
accordingly.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]