jarredhj0214 commented on PR #13119:
URL: https://github.com/apache/gravitino/pull/13119#issuecomment-5748346798

   > Could you provide some details about the error for this scenario
   
   > Could you provide some details about the error for this scenario?
   
   Sure. Let me provide more details.
   
   This PR is not intended to use a switch to hide a specific bug. The real 
problem is about platform-wide adoption and migration control.
   
   In our production environment, `spark.plugins`, `spark.sql.extensions`, and 
some catalog configurations are often injected by a platform-level Spark 
template, instead of being manually configured by each individual Spark job. 
For example:
   
   ```properties
   
spark.plugins=org.apache.gravitino.spark.connector.plugin.GravitinoSparkPlugin
   spark.sql.gravitino.uri=...
   spark.sql.gravitino.metalake=...
   spark.sql.catalog.xxx=...
   spark.sql.extensions=...
   ```
   
   Once Gravitino is enabled in the common template, many existing Spark jobs 
will load Gravitino-related components even if they do not really intend to use 
Gravitino.
   
   We have two typical kinds of cases.
   
   The first one is an unsupported component that we do not plan to support, 
such as some existing TiSpark jobs. These jobs need a gradual migration or 
deprecation period, but they should not block Gravitino adoption for all other 
jobs. One example error is:
   
   ```text
   java.lang.UnsupportedOperationException: class 
com.pingcap.tispark.v2.TiDBTable$$Lambda$5977/764654254: Batch scan are not 
supported
     at org.apache.spark.sql.connector.read.Scan.toBatch(Scan.java:76)
     at 
org.apache.spark.sql.execution.datasources.v2.BatchScanExec.batch$lzycompute(BatchScanExec.scala:42)
     at 
org.apache.spark.sql.execution.datasources.v2.BatchScanExec.batch(BatchScanExec.scala:42)
     at 
org.apache.spark.sql.execution.datasources.v2.BatchScanExec.inputPartitions$lzycompute(BatchScanExec.scala:54)
   ```
   
   The second one is a compatibility issue that can be fixed. For example, we 
observed a Paimon-related compatibility issue and fixed it later. But before 
the fix is available and fully rolled out, affected jobs still need a way to 
opt out temporarily:
   
   ```text
   java.lang.UnsupportedOperationException: Not support TimeType{precision=0}
     at org.apache.gravitino.spark.connector.SparkTypeConverter.toSparkType(...)
   ```
   
   These two cases are different. One may require migration or deprecation, and 
the other can be fixed by improving compatibility. But neither of them should 
block the overall Gravitino rollout.
   
   Without a job-level switch, platform adoption becomes an all-or-nothing 
choice:
   
   - either enable Gravitino globally and risk breaking existing jobs that are 
not ready for it;
   - or disable Gravitino globally and block gradual adoption for jobs that can 
already use it.
   
   The proposed `spark.sql.gravitino.enabled=false` provides a middle path. It 
allows the platform to keep the global Gravitino template, while letting 
specific jobs opt out of Gravitino behavior until they are migrated, 
deprecated, or the corresponding compatibility issues are resolved.
   
   We have already supported this switch in our local downstream repository. I 
raised this PR because I believe this is a common problem that platform users 
may face when adopting Gravitino, so I suggest the community consider accepting 
it.
   
   So the value of this switch is not only for one specific component such as 
TiSpark or Paimon. It is mainly to make Gravitino safer to adopt gradually in 
real production platforms.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to