jarredhj0214 commented on PR #13119:
URL: https://github.com/apache/gravitino/pull/13119#issuecomment-5748346798
> Could you provide some details about the error for this scenario
> Could you provide some details about the error for this scenario?
Sure. Let me provide more details.
This PR is not intended to use a switch to hide a specific bug. The real
problem is about platform-wide adoption and migration control.
In our production environment, `spark.plugins`, `spark.sql.extensions`, and
some catalog configurations are often injected by a platform-level Spark
template, instead of being manually configured by each individual Spark job.
For example:
```properties
spark.plugins=org.apache.gravitino.spark.connector.plugin.GravitinoSparkPlugin
spark.sql.gravitino.uri=...
spark.sql.gravitino.metalake=...
spark.sql.catalog.xxx=...
spark.sql.extensions=...
```
Once Gravitino is enabled in the common template, many existing Spark jobs
will load Gravitino-related components even if they do not really intend to use
Gravitino.
We have two typical kinds of cases.
The first one is an unsupported component that we do not plan to support,
such as some existing TiSpark jobs. These jobs need a gradual migration or
deprecation period, but they should not block Gravitino adoption for all other
jobs. One example error is:
```text
java.lang.UnsupportedOperationException: class
com.pingcap.tispark.v2.TiDBTable$$Lambda$5977/764654254: Batch scan are not
supported
at org.apache.spark.sql.connector.read.Scan.toBatch(Scan.java:76)
at
org.apache.spark.sql.execution.datasources.v2.BatchScanExec.batch$lzycompute(BatchScanExec.scala:42)
at
org.apache.spark.sql.execution.datasources.v2.BatchScanExec.batch(BatchScanExec.scala:42)
at
org.apache.spark.sql.execution.datasources.v2.BatchScanExec.inputPartitions$lzycompute(BatchScanExec.scala:54)
```
The second one is a compatibility issue that can be fixed. For example, we
observed a Paimon-related compatibility issue and fixed it later. But before
the fix is available and fully rolled out, affected jobs still need a way to
opt out temporarily:
```text
java.lang.UnsupportedOperationException: Not support TimeType{precision=0}
at org.apache.gravitino.spark.connector.SparkTypeConverter.toSparkType(...)
```
These two cases are different. One may require migration or deprecation, and
the other can be fixed by improving compatibility. But neither of them should
block the overall Gravitino rollout.
Without a job-level switch, platform adoption becomes an all-or-nothing
choice:
- either enable Gravitino globally and risk breaking existing jobs that are
not ready for it;
- or disable Gravitino globally and block gradual adoption for jobs that can
already use it.
The proposed `spark.sql.gravitino.enabled=false` provides a middle path. It
allows the platform to keep the global Gravitino template, while letting
specific jobs opt out of Gravitino behavior until they are migrated,
deprecated, or the corresponding compatibility issues are resolved.
We have already supported this switch in our local downstream repository. I
raised this PR because I believe this is a common problem that platform users
may face when adopting Gravitino, so I suggest the community consider accepting
it.
So the value of this switch is not only for one specific component such as
TiSpark or Paimon. It is mainly to make Gravitino safer to adopt gradually in
real production platforms.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]