This is an automated email from the ASF dual-hosted git repository.
morningman pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/doris.git
The following commit(s) were added to refs/heads/master by this push:
new 01463f22bc7 [test](fe) Run the ADBC connector unit tests in FE UT
without the arrow-c-data JNI shim (#68510)
01463f22bc7 is described below
commit 01463f22bc715f423b22719cab3ba3dcf384699b
Author: Mingyu Chen (Rayner) <[email protected]>
AuthorDate: Fri Sep 25 17:43:20 2026 +0800
[test](fe) Run the ADBC connector unit tests in FE UT without the
arrow-c-data JNI shim (#68510)
### What problem does this PR solve?
Issue Number: None
Related PR: #66399
Problem Summary:
**Context.** `fe-connector-adbc` is not a Maven prerequisite of fe-core,
so `-am` never reached it and its tests never ran in FE UT; #66399
therefore names the module in the runner. Running it there for the first
time turned 23 of its tests red on the CI agent, all with the same
error:
```
java.lang.UnsatisfiedLinkError: /tmp/jnilib-*.tmp: /lib64/libstdc++.so.6:
version `CXXABI_1.3.9' not found (required by /tmp/jnilib-*.tmp)
```
Two different native libraries are on that path, and only one of them
was ours. The ADBC JNI shim is built by `thirdparty` locally, which is
why it loads anywhere. Reading a driver's result as Arrow additionally
needs arrow-c-data's `arrow_cdata_jni`, which travels inside the
arrow-c-data jar as an **upstream** build: it is selected through the
`arrow.cdata.library.path` property, nothing in this repository sets
that property, so the jar's binary is always the one loaded -- and it
needs CXXABI_1.3.9 (GCC 4.9+), which a host Doris still supports does
not have (CentOS 7 stops at CXXABI_1.3.7).
**1. The problem, and what it cost**
- Every ADBC test that materializes Arrow data (23 of the module's 204)
fails FE UT on such a host, with an error that reads like a test bug
rather than a platform limit.
- The rule was nowhere written down, and it is not visible from the
test: a case that iterates an `ArrowReader` looks exactly like one that
does not, so the next person adding a test to this module walks into the
same wall.
- The suites delegate layer-level questions to these unit tests
(`test_adbc_metadata_ops`, `test_adbc_sqlite_catalog_scan`,
`test_adbc_catalog_scan` and `test_adbc_predicate_pushdown` each say so
in a comment), so deleting them without replacements would have moved
that coverage out of CI silently.
**2. What this PR does, and why it helps**
Commit 1 -- `[test](fe) Drop the ADBC tests that need arrow-c-data's JNI
shim, keep the rest in FE UT`:
- The three classes that read a driver's result are deleted:
`AdbcQueryBuilderNativeTest` (its text-level twin `AdbcQueryBuilderTest`
stays, and "the source accepts the generated SQL" is now asserted end to
end), `AdbcMetadataCacheNativeTest` (the cache's mechanics stay covered
by the pure-Java `AdbcMetadataCacheTest`, 18 cases), and
`AdbcConnectorMetadataTest`.
- Of that last class, the two cases that need a driver but never
materialize a result -- a missing table must not be reported as a driver
gap, and the table descriptor must be typed for the scan path -- move to
`AdbcConnectorMetadataNativeTest` unchanged.
- `AdbcObjectsReaderTest` is rewritten against hand-built `getObjects`
responses (the file already carried the in-memory Arrow technique), so
its four parsing cases run in FE UT with no native library at all; two
more cases are added for the shapes a driver reports for an absent
schema level, and for a namespace reported twice.
- `AdbcNativeTestSupport` now states the boundary: tests that iterate an
`ArrowReader` or touch `org.apache.arrow.c` belong in
`regression-test/suites/external_table_p0/adbc`, not in this module's
test sources.
Commit 2 -- `[test](regression) Cover the ADBC cases the removed unit
tests pinned`: the end-to-end half of those behaviours, as assertions
rather than baselines, in the four suites listed below.
**What it buys:** FE UT runs 187 ADBC tests with no dependency on the
upstream C-data binary; the behaviours that need a real driver or a real
cluster are asserted where they can be; and the rule lives in the class
that owns the boundary.
**3. The classes, and how they call each other**
- `AdbcObjectsReader` -- parses the nested `getObjects` response. Now
fed by fixtures that reproduce the same shapes a driver emits (a schema
entry named the empty string, a null schema list, advisory filters that
return other namespaces' rows), and covered again in FE UT.
- `AdbcNativeTestSupport` -- the single funnel every driver-backed ADBC
test goes through; its javadoc is where the two-library distinction is
now recorded.
- `AdbcConnectorMetadata` / `AdbcMetadataCache` -- the two surviving
driver-backed cases never reach a result (one throws before it, one
never asks the source), which is what makes them safe in FE UT.
- Regression suites -- `test_adbc_sqlite_catalog_scan` (empty-named
database, BLOB mapping, schema served until REFRESH),
`test_adbc_metadata_ops` (same REFRESH negative),
`test_adbc_catalog_scan` (a second database in the source; the sibling
listing asserted as an exact set), `test_adbc_predicate_pushdown` (the
remote statement's two-level qualification).
- Not touched: no FE source change, and `run-fe-ut.sh` keeps the module
line, which is the point of the whole change.
```
run-fe-ut.sh ──▶ fe-connector-adbc (187 tests)
│
├─ driver-backed, reads no result ──▶
AdbcNativeTestSupport ──▶ thirdparty-built libadbc_driver_jni.so
│ (skips when
absent; never touches arrow-c-data)
└─ pure Java ──▶ AdbcObjectsReaderTest fixtures
(in-memory Arrow IPC)
▲
regression suites (real driver, real cluster) ──┘ same shapes, asserted
end to end
```
---
.../adbc/AdbcConnectorMetadataNativeTest.java | 107 +++++++
.../connector/adbc/AdbcConnectorMetadataTest.java | 221 ---------------
.../adbc/AdbcMetadataCacheNativeTest.java | 184 -------------
.../connector/adbc/AdbcNativeTestSupport.java | 9 +
.../connector/adbc/AdbcObjectsReaderTest.java | 306 +++++++++++++++------
.../connector/adbc/AdbcQueryBuilderNativeTest.java | 212 --------------
.../adbc/test_adbc_catalog_scan.groovy | 22 ++
.../adbc/test_adbc_metadata_ops.groovy | 9 +
.../adbc/test_adbc_predicate_pushdown.groovy | 8 +
.../adbc/test_adbc_sqlite_catalog_scan.groovy | 27 +-
run-fe-ut.sh | 3 +
11 files changed, 403 insertions(+), 705 deletions(-)
diff --git
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcConnectorMetadataNativeTest.java
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcConnectorMetadataNativeTest.java
new file mode 100644
index 00000000000..86110844f55
--- /dev/null
+++
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcConnectorMetadataNativeTest.java
@@ -0,0 +1,107 @@
+// Licensed to the Apache Software Foundation (ASF) under one
+// or more contributor license agreements. See the NOTICE file
+// distributed with this work for additional information
+// regarding copyright ownership. The ASF licenses this file
+// to you under the Apache License, Version 2.0 (the
+// "License"); you may not use this file except in compliance
+// with the License. You may obtain a copy of the License at
+//
+// http://www.apache.org/licenses/LICENSE-2.0
+//
+// Unless required by applicable law or agreed to in writing,
+// software distributed under the License is distributed on an
+// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
+// KIND, either express or implied. See the License for the
+// specific language governing permissions and limitations
+// under the License.
+
+package org.apache.doris.connector.adbc;
+
+import org.apache.doris.connector.spi.DorisConnectorException;
+import org.apache.doris.thrift.TTableDescriptor;
+import org.apache.doris.thrift.TTableType;
+
+import org.apache.arrow.adbc.core.AdbcStatement;
+import org.junit.jupiter.api.Assertions;
+import org.junit.jupiter.api.Test;
+import org.junit.jupiter.api.io.TempDir;
+
+import java.nio.file.Path;
+import java.util.Map;
+
+/**
+ * The two metadata cases that need a real driver but never read a result: a
missing table must not be
+ * reported as a driver gap, and the table descriptor handed to the scan path
must be typed.
+ *
+ * <p>Both stop before any Arrow data is materialized -- the first ends in a
thrown exception, the second
+ * never asks the source -- which is what makes them safe to run in FE UT. See
{@link AdbcNativeTestSupport}
+ * for the rule: a test that iterates an {@code ArrowReader} also needs
arrow-c-data's JNI shim, which is an
+ * upstream binary that cannot load on every host Doris supports, and those
tests belong in the regression
+ * suites instead. The rest of the metadata surface -- listings, schema
mapping, views, handles -- is
+ * asserted end to end by {@code
regression-test/suites/external_table_p0/adbc}, which is where it can be.
+ */
+class AdbcConnectorMetadataNativeTest {
+
+ private static AdbcClient sqliteClient(Path dbFile) {
+ return new AdbcClient(AdbcNativeTestSupport.sqliteDriver(),
"libadbc_driver_sqlite.so",
+ null, "file:" + dbFile, null, null, Map.of());
+ }
+
+ /**
+ * A fresh cache per call, so each test reads the source rather than an
earlier test's answers.
+ */
+ private static AdbcConnectorMetadata metadataOn(AdbcClient client) {
+ return new AdbcConnectorMetadata(client, new AdbcSchemaStrategy(),
+ AdbcDialectRegistry::defaultDialect, new
AdbcMetadataCache(Map.of()));
+ }
+
+ private static void seed(AdbcClient client) {
+ client.withConnection(connection -> {
+ for (String sql : new String[] {
+ "CREATE TABLE IF NOT EXISTS t1 (c_int INTEGER, c_dbl REAL,
c_txt TEXT, c_blob BLOB)",
+ "INSERT INTO t1 VALUES (1, 1.5, 'a', x'00ff')",
+ "CREATE TABLE IF NOT EXISTS t2 (a INTEGER)",
+ "CREATE VIEW IF NOT EXISTS v1 AS SELECT * FROM t1"}) {
+ try (AdbcStatement statement = connection.createStatement()) {
+ statement.setSqlQuery(sql);
+ statement.executeUpdate();
+ }
+ }
+ return null;
+ });
+ }
+
+ @Test
+ void missingTableIsReportedAsSuchNotAsADriverGap(@TempDir Path tempDir) {
+ try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
+ seed(client);
+ AdbcConnectorMetadata metadata = metadataOn(client);
+ // Build a handle for a table that does not exist, bypassing
getTableHandle's existence check.
+ AdbcTableHandle ghost = new AdbcTableHandle(new
AdbcNamespace("main", ""), "no_such_table");
+
+ DorisConnectorException e =
Assertions.assertThrows(DorisConnectorException.class,
+ () -> metadata.getTableSchema(null, ghost));
+
+ // The fallback to executeSchema must fire only on
NOT_IMPLEMENTED. Falling back on every error
+ // would answer a plain missing table with "this driver implements
neither method", sending the
+ // user to look at their driver instead of their table name.
+ Assertions.assertTrue(e.getMessage().contains("no_such_table"),
e.getMessage());
+ Assertions.assertFalse(e.getMessage().contains("implements
neither"), e.getMessage());
+ }
+ }
+
+ @Test
+ void tableDescriptorIsTypedForTheScanPath(@TempDir Path tempDir) {
+ try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
+ seed(client);
+
+ TTableDescriptor descriptor = metadataOn(client)
+ .buildTableDescriptor(null, 7L, "t1", "main", "t1", 4,
42L);
+
+ // Returning null (the SPI default) would let fe-core fall back to
SCHEMA_TABLE, and BE would
+ // then build a SchemaTableDescriptor rather than the one the
file-scan path expects.
+ Assertions.assertEquals(TTableType.HIVE_TABLE,
descriptor.getTableType());
+ Assertions.assertTrue(descriptor.isSetHiveTable());
+ }
+ }
+}
diff --git
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcConnectorMetadataTest.java
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcConnectorMetadataTest.java
deleted file mode 100644
index cd2dc5ff9c3..00000000000
---
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcConnectorMetadataTest.java
+++ /dev/null
@@ -1,221 +0,0 @@
-// Licensed to the Apache Software Foundation (ASF) under one
-// or more contributor license agreements. See the NOTICE file
-// distributed with this work for additional information
-// regarding copyright ownership. The ASF licenses this file
-// to you under the Apache License, Version 2.0 (the
-// "License"); you may not use this file except in compliance
-// with the License. You may obtain a copy of the License at
-//
-// http://www.apache.org/licenses/LICENSE-2.0
-//
-// Unless required by applicable law or agreed to in writing,
-// software distributed under the License is distributed on an
-// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-// KIND, either express or implied. See the License for the
-// specific language governing permissions and limitations
-// under the License.
-
-package org.apache.doris.connector.adbc;
-
-import org.apache.doris.connector.spi.ConnectorColumn;
-import org.apache.doris.connector.spi.ConnectorTableSchema;
-import org.apache.doris.connector.spi.ConnectorType;
-import org.apache.doris.connector.spi.DorisConnectorException;
-import org.apache.doris.connector.spi.handle.ConnectorColumnHandle;
-import org.apache.doris.connector.spi.handle.ConnectorTableHandle;
-import org.apache.doris.thrift.TTableDescriptor;
-import org.apache.doris.thrift.TTableType;
-
-import org.apache.arrow.adbc.core.AdbcStatement;
-import org.junit.jupiter.api.Assertions;
-import org.junit.jupiter.api.Test;
-import org.junit.jupiter.api.io.TempDir;
-
-import java.nio.file.Path;
-import java.util.ArrayList;
-import java.util.List;
-import java.util.Map;
-import java.util.Optional;
-
-/**
- * The metadata surface end to end against the real SQLite ADBC driver: the
same Java -> JNI -> C driver
- * manager -> driver .so path FE takes in production, with only the source
swapped for a temporary file.
- *
- * <p>Skips loudly when thirdparty's native libraries are absent -- see {@link
AdbcNativeTestSupport}. A
- * skipped run has verified nothing here.
- */
-class AdbcConnectorMetadataTest {
-
- private static AdbcClient sqliteClient(Path dbFile) {
- return new AdbcClient(AdbcNativeTestSupport.sqliteDriver(),
"libadbc_driver_sqlite.so",
- null, "file:" + dbFile, null, null, Map.of());
- }
-
- /**
- * A fresh cache per call, so each test here reads the source rather than
the previous test's answers --
- * what the cache does across statements is {@link
AdbcMetadataCacheNativeTest}'s subject, not this one's.
- */
- private static AdbcConnectorMetadata metadataOn(AdbcClient client) {
- return new AdbcConnectorMetadata(client, new AdbcSchemaStrategy(),
- AdbcDialectRegistry::defaultDialect, new
AdbcMetadataCache(Map.of()));
- }
-
- /**
- * The INSERT is load-bearing, not decoration. SQLite is dynamically typed
and its ADBC driver derives
- * the Arrow schema from the values present, not from the declared column
types: on an empty t1 it
- * reports all four columns as int64. So a fixture with no rows would
assert a type mapping the driver
- * never actually produced.
- */
- private static void seed(AdbcClient client) {
- client.withConnection(connection -> {
- for (String sql : new String[] {
- "CREATE TABLE IF NOT EXISTS t1 (c_int INTEGER, c_dbl REAL,
c_txt TEXT, c_blob BLOB)",
- "INSERT INTO t1 VALUES (1, 1.5, 'a', x'00ff')",
- "CREATE TABLE IF NOT EXISTS t2 (a INTEGER)",
- "CREATE VIEW IF NOT EXISTS v1 AS SELECT * FROM t1"}) {
- try (AdbcStatement statement = connection.createStatement()) {
- statement.setSqlQuery(sql);
- statement.executeUpdate();
- }
- }
- return null;
- });
- }
-
- @Test
- void showDatabasesReturnsTheFlattenedNamespace(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
- AdbcConnectorMetadata metadata = metadataOn(client);
-
- Assertions.assertEquals(List.of("main"),
metadata.listDatabaseNames(null));
- Assertions.assertTrue(metadata.databaseExists(null, "main"));
- Assertions.assertFalse(metadata.databaseExists(null,
"no_such_db"));
- }
- }
-
- @Test
- void showTablesExcludesViews(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
-
- // Doris presents no views for an ADBC catalog, so listing v1
would produce a name that DESC and
- // SELECT would both then fail on.
- Assertions.assertEquals(List.of("t1", "t2"),
metadataOn(client).listTableNames(null, "main"));
- }
- }
-
- @Test
- void listingTablesOfAnUnknownDatabaseIsEmptyNotAnError(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
- Assertions.assertEquals(List.of(),
metadataOn(client).listTableNames(null, "no_such_db"));
- }
- }
-
- @Test
- void tableHandleCarriesTheRemoteCoordinates(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
-
- Optional<ConnectorTableHandle> handle =
metadataOn(client).getTableHandle(null, "main", "t1");
-
- Assertions.assertTrue(handle.isPresent());
- AdbcTableHandle adbc = (AdbcTableHandle) handle.get();
- // SQLite names a catalog but no schema, so the Doris database
name came from the catalog level.
- // The handle must still record which level it came from, or the
pushed-down SQL later cannot
- // qualify the table.
- Assertions.assertEquals("main", adbc.getRemoteCatalog());
- Assertions.assertEquals("", adbc.getRemoteDbSchema());
- Assertions.assertEquals("t1", adbc.getRemoteTable());
- Assertions.assertEquals("main", adbc.getDorisDbName());
- }
- }
-
- @Test
- void unknownTableAndUnknownDatabaseBothYieldNoHandle(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
- AdbcConnectorMetadata metadata = metadataOn(client);
-
- Assertions.assertEquals(Optional.empty(),
metadata.getTableHandle(null, "main", "no_such"));
- Assertions.assertEquals(Optional.empty(),
metadata.getTableHandle(null, "no_such", "t1"));
- // A view is not a table here; handing back a handle for it would
defer the failure to DESC.
- Assertions.assertEquals(Optional.empty(),
metadata.getTableHandle(null, "main", "v1"));
- }
- }
-
- @Test
- void descMapsTheRealArrowSchema(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
- AdbcConnectorMetadata metadata = metadataOn(client);
- ConnectorTableHandle handle = metadata.getTableHandle(null,
"main", "t1").orElseThrow();
-
- ConnectorTableSchema schema = metadata.getTableSchema(null,
handle);
-
- // Expected values come from what the driver actually reports:
SQLite INTEGER arrives as int64,
- // REAL as float64, TEXT as utf8, BLOB as binary. Doris has no
binary column type here, so BLOB
- // lands on STRING.
- List<String> names = new ArrayList<>();
- List<ConnectorType> types = new ArrayList<>();
- for (ConnectorColumn column : schema.getColumns()) {
- names.add(column.getName());
- types.add(column.getType());
- }
- Assertions.assertEquals(List.of("c_int", "c_dbl", "c_txt",
"c_blob"), names);
- Assertions.assertEquals(List.of(ConnectorType.of("BIGINT"),
ConnectorType.of("DOUBLE"),
- ConnectorType.of("STRING"), ConnectorType.of("STRING")),
types);
- Assertions.assertEquals("t1", schema.getTableName());
- }
- }
-
- @Test
- void columnHandlesMatchTheSchemaAndItsOrder(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
- AdbcConnectorMetadata metadata = metadataOn(client);
- ConnectorTableHandle handle = metadata.getTableHandle(null,
"main", "t1").orElseThrow();
-
- Map<String, ConnectorColumnHandle> handles =
metadata.getColumnHandles(null, handle);
-
- // Order matters: the scan path pairs handles with schema columns
positionally.
- Assertions.assertEquals(List.of("c_int", "c_dbl", "c_txt",
"c_blob"),
- new ArrayList<>(handles.keySet()));
- }
- }
-
- @Test
- void missingTableIsReportedAsSuchNotAsADriverGap(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
- AdbcConnectorMetadata metadata = metadataOn(client);
- // Build a handle for a table that does not exist, bypassing
getTableHandle's existence check.
- AdbcTableHandle ghost = new AdbcTableHandle(new
AdbcNamespace("main", ""), "no_such_table");
-
- DorisConnectorException e =
Assertions.assertThrows(DorisConnectorException.class,
- () -> metadata.getTableSchema(null, ghost));
-
- // The fallback to executeSchema must fire only on
NOT_IMPLEMENTED. Falling back on every error
- // would answer a plain missing table with "this driver implements
neither method", sending the
- // user to look at their driver instead of their table name.
- Assertions.assertTrue(e.getMessage().contains("no_such_table"),
e.getMessage());
- Assertions.assertFalse(e.getMessage().contains("implements
neither"), e.getMessage());
- }
- }
-
- @Test
- void tableDescriptorIsTypedForTheScanPath(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("meta.db"))) {
- seed(client);
-
- TTableDescriptor descriptor = metadataOn(client)
- .buildTableDescriptor(null, 7L, "t1", "main", "t1", 4,
42L);
-
- // Returning null (the SPI default) would let fe-core fall back to
SCHEMA_TABLE, and BE would
- // then build a SchemaTableDescriptor rather than the one the
file-scan path expects.
- Assertions.assertEquals(TTableType.HIVE_TABLE,
descriptor.getTableType());
- Assertions.assertTrue(descriptor.isSetHiveTable());
- }
- }
-}
diff --git
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcMetadataCacheNativeTest.java
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcMetadataCacheNativeTest.java
deleted file mode 100644
index 212c008651c..00000000000
---
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcMetadataCacheNativeTest.java
+++ /dev/null
@@ -1,184 +0,0 @@
-// Licensed to the Apache Software Foundation (ASF) under one
-// or more contributor license agreements. See the NOTICE file
-// distributed with this work for additional information
-// regarding copyright ownership. The ASF licenses this file
-// to you under the Apache License, Version 2.0 (the
-// "License"); you may not use this file except in compliance
-// with the License. You may obtain a copy of the License at
-//
-// http://www.apache.org/licenses/LICENSE-2.0
-//
-// Unless required by applicable law or agreed to in writing,
-// software distributed under the License is distributed on an
-// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-// KIND, either express or implied. See the License for the
-// specific language governing permissions and limitations
-// under the License.
-
-package org.apache.doris.connector.adbc;
-
-import org.apache.doris.connector.spi.ConnectorColumn;
-import org.apache.doris.connector.spi.ConnectorTableSchema;
-import org.apache.doris.connector.spi.handle.ConnectorTableHandle;
-
-import org.apache.arrow.adbc.core.AdbcStatement;
-import org.junit.jupiter.api.Assertions;
-import org.junit.jupiter.api.Test;
-import org.junit.jupiter.api.io.TempDir;
-
-import java.nio.file.Path;
-import java.util.ArrayList;
-import java.util.List;
-import java.util.Map;
-import java.util.Optional;
-
-/**
- * The metadata path with a catalog-level cache in front of it, against the
real SQLite driver.
- *
- * <p>Each test changes the source behind Doris's back and then asks what
Doris sees. That is the only
- * evidence that says whether an answer came from memory or from the driver,
and unlike a call counter it
- * cannot be satisfied by a cache that stores things it never reads.
- *
- * <p>Every {@code metadata()} call stands for one statement: the engine
builds a fresh
- * {@link AdbcConnectorMetadata} per statement, and the cache is what they
share.
- *
- * <p>Skips loudly when thirdparty's native libraries are absent -- see {@link
AdbcNativeTestSupport}.
- */
-class AdbcMetadataCacheNativeTest {
-
- private final AdbcMetadataCache cache = new AdbcMetadataCache(Map.of());
-
- private static AdbcClient sqliteClient(Path dbFile) {
- return new AdbcClient(AdbcNativeTestSupport.sqliteDriver(),
"libadbc_driver_sqlite.so",
- null, "file:" + dbFile, null, null, Map.of());
- }
-
- /** One statement's view of the catalog. Separate objects, one shared
cache -- as in production. */
- private AdbcConnectorMetadata metadata(AdbcClient client) {
- return new AdbcConnectorMetadata(client, new AdbcSchemaStrategy(),
- AdbcDialectRegistry::defaultDialect, cache);
- }
-
- private static void execute(AdbcClient client, String... statements) {
- client.withConnection(connection -> {
- for (String sql : statements) {
- try (AdbcStatement statement = connection.createStatement()) {
- statement.setSqlQuery(sql);
- statement.executeUpdate();
- }
- }
- return null;
- });
- }
-
- /** SQLite derives its Arrow types from the values present, so a row is
needed for the types to be real. */
- private static void seed(AdbcClient client) {
- execute(client,
- "CREATE TABLE t1 (c_int INTEGER, c_txt TEXT)",
- "INSERT INTO t1 VALUES (1, 'a')");
- }
-
- private static List<String> columnNames(ConnectorTableSchema schema) {
- List<String> names = new ArrayList<>();
- for (ConnectorColumn column : schema.getColumns()) {
- names.add(column.getName());
- }
- return names;
- }
-
- private List<String> columnsOf(AdbcClient client, String table) {
- ConnectorTableHandle handle = metadata(client).getTableHandle(null,
"main", table).orElseThrow();
- return columnNames(metadata(client).getTableSchema(null, handle));
- }
-
- @Test
- void theNextStatementReadsTheSchemaTheLastOneAlreadyPaidFor(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("cache.db"))) {
- seed(client);
- Assertions.assertEquals(List.of("c_int", "c_txt"),
columnsOf(client, "t1"));
-
- execute(client, "ALTER TABLE t1 ADD COLUMN c_added INTEGER");
-
- // The column really is there now -- the source changed and Doris
was not told. Serving the
- // remembered shape is the whole point; noticing the change here
would mean nothing was cached.
- Assertions.assertEquals(List.of("c_int", "c_txt"),
columnsOf(client, "t1"));
- }
- }
-
- @Test
- void refreshTableIsWhatMakesTheAlteredColumnsVisible(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("cache.db"))) {
- seed(client);
- columnsOf(client, "t1");
- execute(client, "ALTER TABLE t1 ADD COLUMN c_added INTEGER");
-
- cache.invalidateTable("main", "t1");
-
- Assertions.assertEquals(List.of("c_int", "c_txt", "c_added"),
columnsOf(client, "t1"));
- }
- }
-
- /**
- * Decision C. Reading the listing from memory is fine; concluding from
memory that a name does not exist
- * is not. A user who just created a table and is told it is not there has
no way to tell that from a
- * typo, and no reason to suspect a cache.
- */
- @Test
- void tableCreatedAfterTheListingWasCachedIsStillFound(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("cache.db"))) {
- seed(client);
- metadata(client).listTableNames(null, "main");
-
- execute(client, "CREATE TABLE t_new (a INTEGER)", "INSERT INTO
t_new VALUES (1)");
-
- Optional<ConnectorTableHandle> handle =
metadata(client).getTableHandle(null, "main", "t_new");
- Assertions.assertTrue(handle.isPresent(), "a table created after
the listing was cached must"
- + " still be reachable by name");
- Assertions.assertEquals(List.of("a"), columnNames(metadata(client)
- .getTableSchema(null, handle.get())));
- }
- }
-
- @Test
- void missingTableIsStillMissingAfterTheListingIsReRead(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("cache.db"))) {
- seed(client);
- metadata(client).listTableNames(null, "main");
-
- // Re-reading the listing is a last chance to find the name, not a
way to accept any name.
- Assertions.assertEquals(Optional.empty(),
- metadata(client).getTableHandle(null, "main",
"no_such_table"));
- }
- }
-
- /**
- * The listing methods stay live however much is remembered. They read
like reports, but the engine loads
- * its own name cache from them and then decides from that whether a table
exists at all -- including the
- * re-list it does as a last chance for a name it has never seen. A cached
answer here would turn that
- * re-check into a formality and leave a table created a moment ago
unreachable.
- */
- @Test
- void listingTheTablesAlwaysAsksTheSource(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("cache.db"))) {
- seed(client);
- Assertions.assertEquals(List.of("t1"),
metadata(client).listTableNames(null, "main"));
-
- execute(client, "CREATE TABLE t_new (a INTEGER)");
-
- Assertions.assertEquals(List.of("t1", "t_new"),
metadata(client).listTableNames(null, "main"));
- }
- }
-
- @Test
- void listingTheDatabasesAlwaysAsksTheSource(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("cache.db"))) {
- seed(client);
- // Plant a database the source does not have. Creating one behind
Doris's back is what this
- // stands in for -- SQLite gains a catalog only by ATTACH, which
does not outlive the connection
- // it ran on -- and it fails the same way: anything answered from
memory shows the ghost.
- cache.namespaces(() -> List.of(new AdbcNamespace("ghost", "")));
-
- Assertions.assertEquals(List.of("main"),
metadata(client).listDatabaseNames(null));
- }
- }
-}
diff --git
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcNativeTestSupport.java
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcNativeTestSupport.java
index 72d05bd4478..7b36ad59ce9 100644
---
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcNativeTestSupport.java
+++
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcNativeTestSupport.java
@@ -36,6 +36,15 @@ import java.nio.file.Paths;
*
* <p>When the libraries are absent the tests SKIP, and say so loudly: a
skipped run has verified nothing
* about the native path and must not be read as a pass.
+ *
+ * <p><b>These libraries are not the whole native surface.</b> A test that
reads a driver's result into
+ * Arrow also loads arrow-c-data's own JNI shim, and that one is not built
here: it travels inside the
+ * arrow-c-data jar as an upstream build that needs CXXABI_1.3.9, which hosts
Doris still supports do not
+ * have (CentOS 7's libstdc++ stops at CXXABI_1.3.7). Nothing in this
repository redirects it, so a test
+ * that materializes Arrow data does not exercise the connector on such a host
-- it fails FE UT, which is
+ * how this rule was learned. Tests that iterate an {@code ArrowReader}, or
touch {@code org.apache.arrow.c},
+ * therefore belong in the regression suites under {@code
regression-test/suites/external_table_p0/adbc},
+ * not in this module's test sources.
*/
final class AdbcNativeTestSupport {
diff --git
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcObjectsReaderTest.java
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcObjectsReaderTest.java
index 793c9327a61..92da09f5f62 100644
---
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcObjectsReaderTest.java
+++
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcObjectsReaderTest.java
@@ -19,12 +19,12 @@ package org.apache.doris.connector.adbc;
import org.apache.doris.connector.spi.DorisConnectorException;
-import org.apache.arrow.adbc.core.AdbcConnection;
-import org.apache.arrow.adbc.core.AdbcStatement;
import org.apache.arrow.memory.BufferAllocator;
import org.apache.arrow.memory.RootAllocator;
import org.apache.arrow.vector.VarCharVector;
import org.apache.arrow.vector.VectorSchemaRoot;
+import org.apache.arrow.vector.complex.ListVector;
+import org.apache.arrow.vector.complex.StructVector;
import org.apache.arrow.vector.ipc.ArrowReader;
import org.apache.arrow.vector.ipc.ArrowStreamReader;
import org.apache.arrow.vector.ipc.ArrowStreamWriter;
@@ -34,121 +34,158 @@ import org.apache.arrow.vector.types.pojo.FieldType;
import org.apache.arrow.vector.types.pojo.Schema;
import org.junit.jupiter.api.Assertions;
import org.junit.jupiter.api.Test;
-import org.junit.jupiter.api.io.TempDir;
import java.io.ByteArrayInputStream;
import java.io.ByteArrayOutputStream;
import java.nio.channels.Channels;
-import java.nio.file.Path;
+import java.nio.charset.StandardCharsets;
import java.util.List;
-import java.util.Map;
/**
- * Reads {@code getObjects} results produced by the real SQLite driver, so the
nested-Arrow parsing is
- * checked against a shape a driver actually emits rather than one this test
invented. A hand-built result
- * covers the case a real driver will not produce on demand: a non-standard
schema.
+ * {@link AdbcObjectsReader} against hand-built {@code getObjects} responses.
+ *
+ * <p>These cases were once driven through the real SQLite driver, and cannot
be any more. Reading a driver's
+ * result into Arrow also loads arrow-c-data's own JNI shim, which travels
inside the arrow-c-data jar as an
+ * upstream build: it needs CXXABI_1.3.9, and hosts Doris still supports do
not have it (CentOS 7's
+ * libstdc++ stops at CXXABI_1.3.7). No property in this repository redirects
it, so a driver-backed reader
+ * test does not test the reader on such a host -- it fails FE UT. The
response shape is what this class is
+ * about, and that shape is fixed by the ADBC standard; a fixture expresses it
without the driver.
+ *
+ * <p>Two shapes below are worth knowing came from a live driver rather than
from the standard, because a
+ * source that differed would be a real finding: SQLite reports a namespace
whose {@code db_schema_name} is
+ * the empty string (not null, and not a missing entry), and a depth-CATALOGS
answer reports the schema list
+ * as null. The end-to-end half of that -- that a real driver really answers
this way, and that a view stays
+ * out of {@code SHOW TABLES} when the source ignores the type filter -- is
pinned by the regression suites
+ * under {@code regression-test/suites/external_table_p0/adbc}, which have the
driver and the cluster.
*/
class AdbcObjectsReaderTest {
- private static AdbcClient sqliteClient(Path dbFile) {
- return new AdbcClient(AdbcNativeTestSupport.sqliteDriver(),
"libadbc_driver_sqlite.so",
- null, "file:" + dbFile, null, null, Map.of());
- }
+ private static final String CATALOG_NAME = "catalog_name";
+ private static final String CATALOG_DB_SCHEMAS = "catalog_db_schemas";
- private static void seed(AdbcClient client) {
- client.withConnection(connection -> {
- for (String sql : new String[] {
- "CREATE TABLE IF NOT EXISTS t1 (c_int INTEGER, c_txt
TEXT)",
- "CREATE TABLE IF NOT EXISTS t2 (a INTEGER)",
- "CREATE VIEW IF NOT EXISTS v1 AS SELECT * FROM t1"}) {
- try (AdbcStatement statement = connection.createStatement()) {
- statement.setSqlQuery(sql);
- statement.executeUpdate();
- }
- }
- return null;
- });
+ private static final Field TABLE_STRUCT = new Field("item",
FieldType.nullable(new ArrowType.Struct()),
+ List.of(new Field("table_name",
FieldType.nullable(ArrowType.Utf8.INSTANCE), null),
+ new Field("table_type",
FieldType.nullable(ArrowType.Utf8.INSTANCE), null)));
+
+ private static final Field SCHEMA_STRUCT = new Field("item",
FieldType.nullable(new ArrowType.Struct()),
+ List.of(new Field("db_schema_name",
FieldType.nullable(ArrowType.Utf8.INSTANCE), null),
+ new Field("db_schema_tables", FieldType.nullable(new
ArrowType.List()),
+ List.of(TABLE_STRUCT))));
+
+ private static final Schema GET_OBJECTS_SCHEMA = new Schema(List.of(
+ new Field(CATALOG_NAME,
FieldType.nullable(ArrowType.Utf8.INSTANCE), null),
+ new Field(CATALOG_DB_SCHEMAS, FieldType.nullable(new
ArrowType.List()), List.of(SCHEMA_STRUCT))));
+
+ /**
+ * A source with no schema layer: SQLite reports catalog "main" and
answers the schema entry with the
+ * EMPTY STRING, not with null and not by omitting the entry. Reading that
as an absent level is the
+ * difference between a database named "main" and one named "".
+ */
+ @Test
+ void readsTheNamespaceASourceWithNoSchemaLayerReports() throws Exception {
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator, catalog("main",
schema("", null)))) {
+ List<AdbcNamespace> namespaces =
AdbcObjectsReader.readNamespaces(reader);
+
+ Assertions.assertEquals(List.of(new AdbcNamespace("main", "")),
namespaces);
+ Assertions.assertEquals("main",
namespaces.get(0).dorisDatabaseName());
+ }
}
+ /**
+ * The same source asked at depth CATALOGS answers a null schema list,
which is the other way a driver
+ * says "there is no schema level here". Both must land on the catalog
name as the database.
+ */
@Test
- void readsTheNamespaceASourceWithNoSchemaLayerReports(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("objects.db"))) {
- seed(client);
-
- List<AdbcNamespace> namespaces = client.withConnection(connection
-> {
- try (ArrowReader reader = connection.getObjects(
- AdbcConnection.GetObjectsDepth.DB_SCHEMAS, null, null,
null, null, null)) {
- return AdbcObjectsReader.readNamespaces(reader);
- }
- });
+ void nullSchemaListStillYieldsTheCatalogAsADatabase() throws Exception {
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator, catalog("main",
null))) {
+ List<AdbcNamespace> namespaces =
AdbcObjectsReader.readNamespaces(reader);
- // SQLite has catalogs but no schema layer, and reports the
missing level as an EMPTY STRING
- // rather than null -- treating only null as "absent" would
produce a Doris database named "".
- Assertions.assertEquals(1, namespaces.size(),
namespaces.toString());
- Assertions.assertEquals("main",
namespaces.get(0).getRemoteCatalog());
- Assertions.assertEquals("", namespaces.get(0).getRemoteDbSchema());
+ Assertions.assertEquals(List.of(new AdbcNamespace("main", "")),
namespaces);
Assertions.assertEquals("main",
namespaces.get(0).dorisDatabaseName());
}
}
+ /** A namespace reported twice is one namespace: a source may repeat a
catalog level it shares. */
@Test
- void readsTableNamesAndHonoursTheTableTypeFilter(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("objects.db"))) {
- seed(client);
- AdbcNamespace main = new AdbcNamespace("main", "");
-
- List<String> tables = client.withConnection(connection -> {
- try (ArrowReader reader = connection.getObjects(
- AdbcConnection.GetObjectsDepth.TABLES, null, null,
null,
- new String[] {"table"}, null)) {
- return AdbcObjectsReader.readTableNames(reader, main);
- }
- });
+ void namespaceReportedTwiceIsListedOnce() throws Exception {
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator,
+ catalog("main", schema("", null)),
+ catalog("main", schema("", null)))) {
+ Assertions.assertEquals(List.of(new AdbcNamespace("main", "")),
+ AdbcObjectsReader.readNamespaces(reader));
+ }
+ }
- // v1 is a view. Doris presents no views for ADBC catalogs, so
listing it would produce a table
- // that DESC and SELECT both fail on.
- Assertions.assertEquals(List.of("t1", "t2"), tables);
+ /**
+ * A source that honours the type filter answers with base tables only,
and they come back in the order
+ * the source listed them.
+ */
+ @Test
+ void readsTableNamesAndHonoursTheTableTypeFilter() throws Exception {
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator, catalog("main",
+ schema("", table("t1", "TABLE"), table("t2", "BASE
TABLE"))))) {
+ Assertions.assertEquals(List.of("t1", "t2"),
+ AdbcObjectsReader.readTableNames(reader, new
AdbcNamespace("main", "")));
}
}
+ /**
+ * A Doris source is what makes this the load-bearing case: its Flight SQL
endpoint recognises only the
+ * literal "VIEW" as a type filter and answers everything else --
including the "table" ADBC asks with --
+ * by returning every object it has. The response therefore carries
objects the connector never asked
+ * for, and they have to be dropped by the table_type that came back with
them.
+ *
+ * <p>Dropping an unrecognised type is deliberate, and a view is the
reason: a leaked view scans fine
+ * through ADBC, so nothing ever looks broken -- the catalog just offers
an object DESC and SELECT then
+ * fail on. A missing type is kept, because it says nothing about the
object.
+ */
@Test
- void viewsAreDroppedEvenWhenTheSourceIgnoresTheTypeFilter(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("objects.db"))) {
- seed(client);
- AdbcNamespace main = new AdbcNamespace("main", "");
-
- // Asking with no type filter reproduces, through a real driver,
what a Doris source does to the
- // filter the connector does send: its Flight SQL endpoint
recognises only the literal "VIEW" and
- // answers "table" with everything. The guarantee has to survive
that, so it cannot live in the
- // request -- v1 must be gone because of the table_type that came
back with it.
- List<String> tables = client.withConnection(connection -> {
- try (ArrowReader reader = connection.getObjects(
- AdbcConnection.GetObjectsDepth.TABLES, null, null,
null, null, null)) {
- return AdbcObjectsReader.readTableNames(reader, main);
- }
- });
+ void viewsAreDroppedEvenWhenTheSourceIgnoresTheTypeFilter() throws
Exception {
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator, catalog("main",
schema("",
+ table("t1", "TABLE"), table("v1", "VIEW"), table("v2",
"view"),
+ table("s1", "SYSTEM VIEW"), table("a1", "OLAP"),
table("n1", null),
+ table("e1", ""), table("b1", "BASE TABLE"))))) {
+ List<String> tables = AdbcObjectsReader.readTableNames(reader, new
AdbcNamespace("main", ""));
- Assertions.assertEquals(List.of("t1", "t2"), tables);
+ // t1 first, then the two the source left untyped -- a source that
omits the column stays as
+ // usable as it was before the filter existed -- and the Doris
spelling of a base table.
+ Assertions.assertEquals(List.of("t1", "n1", "e1", "b1"), tables);
}
}
+ /**
+ * {@code getObjects} filters are advisory: a driver may answer a narrower
request with everything it
+ * has. Rows that belong to another namespace must not be listed under the
one that was asked for, or a
+ * table of a neighbouring database appears in this one -- and the
reverse, a namespace that matches
+ * neither level, must list nothing rather than everything.
+ */
@Test
- void tablesOfOtherNamespacesAreNotListed(@TempDir Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("objects.db"))) {
- seed(client);
- // getObjects filters are advisory: a driver may answer with
everything it has. Returning rows
- // that belong to a different namespace would list those tables
under the wrong database.
- AdbcNamespace other = new AdbcNamespace("someother", "public");
-
- List<String> tables = client.withConnection(connection -> {
- try (ArrowReader reader = connection.getObjects(
- AdbcConnection.GetObjectsDepth.TABLES, null, null,
null,
- new String[] {"table"}, null)) {
- return AdbcObjectsReader.readTableNames(reader, other);
- }
- });
+ void tablesOfOtherNamespacesAreNotListed() throws Exception {
+ Object[] row = catalog("main", schema("", table("t1", "TABLE")),
+ schema("public", table("p1", "TABLE")));
+ Object[] otherCatalog = catalog("other", schema("", table("o1",
"TABLE")));
- Assertions.assertEquals(List.of(), tables);
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator, row, otherCatalog)) {
+ Assertions.assertEquals(List.of("t1"),
+ AdbcObjectsReader.readTableNames(reader, new
AdbcNamespace("main", "")));
+ }
+ // A second read: the same rows, asked for the other schema of the
same catalog. Its neighbour's
+ // table must not come along, and neither may the other catalog's.
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator, row, otherCatalog)) {
+ Assertions.assertEquals(List.of("p1"),
+ AdbcObjectsReader.readTableNames(reader, new
AdbcNamespace("main", "public")));
+ }
+ try (BufferAllocator allocator = new RootAllocator();
+ ArrowReader reader = response(allocator, row, otherCatalog)) {
+ Assertions.assertEquals(List.of(),
+ AdbcObjectsReader.readTableNames(reader, new
AdbcNamespace("main", "no_such_schema")));
}
}
@@ -192,13 +229,100 @@ class AdbcObjectsReaderTest {
}
}
+ // ========= fixture =========
+
+ /** One catalog row: the name the source reported, and its schema entries
(null at depth CATALOGS). */
+ private static Object[] catalog(String name, Object[]... schemas) {
+ return new Object[] {name, schemas == null ? null : List.of(schemas)};
+ }
+
+ /** One schema entry: the name the source reported (SQLite's is the empty
string) and its tables. */
+ private static Object[] schema(String name, Object[]... tables) {
+ return new Object[] {name, tables == null ? null : List.of(tables)};
+ }
+
+ /** One table: the name and the {@code table_type} the source reported,
either of which may be absent. */
+ private static Object[] table(String name, String type) {
+ return new Object[] {name, type};
+ }
+
+ /**
+ * Builds the getObjects response these rows describe and returns it as
the reader a driver would have
+ * handed over -- written to an IPC stream and read back, so the reader
sees real Arrow vectors rather
+ * than the fixture's own objects.
+ */
+ private static ArrowReader response(BufferAllocator allocator, Object[]...
rows) throws Exception {
+ VectorSchemaRoot root = VectorSchemaRoot.create(GET_OBJECTS_SCHEMA,
allocator);
+ VarCharVector catalogs = (VarCharVector) root.getVector(CATALOG_NAME);
+ ListVector schemas = (ListVector) root.getVector(CATALOG_DB_SCHEMAS);
+ StructVector schemaStruct = (StructVector) schemas.getDataVector();
+ VarCharVector schemaNames = schemaStruct.getChild("db_schema_name",
VarCharVector.class);
+ ListVector tables = schemaStruct.getChild("db_schema_tables",
ListVector.class);
+ StructVector tableStruct = (StructVector) tables.getDataVector();
+ VarCharVector tableNames = tableStruct.getChild("table_name",
VarCharVector.class);
+ VarCharVector tableTypes = tableStruct.getChild("table_type",
VarCharVector.class);
+
+ int schemaCount = 0;
+ int tableCount = 0;
+ for (int row = 0; row < rows.length; row++) {
+ catalogs.setSafe(row, utf8((String) rows[row][0]));
+ List<Object[]> rowSchemas = asList(rows[row][1]);
+ if (rowSchemas == null) {
+ schemas.setNull(row);
+ continue;
+ }
+ schemas.startNewValue(row);
+ for (Object[] entry : rowSchemas) {
+ schemaNames.setSafe(schemaCount, utf8((String) entry[0]));
+ schemaStruct.setIndexDefined(schemaCount);
+ List<Object[]> entryTables = asList(entry[1]);
+ if (entryTables == null) {
+ tables.setNull(schemaCount);
+ } else {
+ tables.startNewValue(schemaCount);
+ for (Object[] entryTable : entryTables) {
+ tableNames.setSafe(tableCount, utf8((String)
entryTable[0]));
+ if (entryTable[1] == null) {
+ tableTypes.setNull(tableCount);
+ } else {
+ tableTypes.setSafe(tableCount, utf8((String)
entryTable[1]));
+ }
+ tableStruct.setIndexDefined(tableCount);
+ tableCount++;
+ }
+ tables.endValue(schemaCount, entryTables.size());
+ }
+ schemaCount++;
+ }
+ schemas.endValue(row, rowSchemas.size());
+ }
+
+ // Children before parents: a list's setValueCount is what gives its
child its count.
+ tableNames.setValueCount(tableCount);
+ tableTypes.setValueCount(tableCount);
+ schemaNames.setValueCount(schemaCount);
+ tables.setValueCount(schemaCount);
+ schemas.setValueCount(rows.length);
+ catalogs.setValueCount(rows.length);
+ root.setRowCount(rows.length);
+
+ ByteArrayOutputStream bytes = new ByteArrayOutputStream();
+ try (ArrowStreamWriter writer = new ArrowStreamWriter(root, null,
Channels.newChannel(bytes))) {
+ writer.start();
+ writer.writeBatch();
+ writer.end();
+ }
+ root.close();
+ return new ArrowStreamReader(new
ByteArrayInputStream(bytes.toByteArray()), allocator);
+ }
+
private static ArrowReader oneRowReader(BufferAllocator allocator, Schema
schema) throws Exception {
ByteArrayOutputStream bytes = new ByteArrayOutputStream();
try (VectorSchemaRoot root = VectorSchemaRoot.create(schema,
allocator);
ArrowStreamWriter writer = new ArrowStreamWriter(root, null,
Channels.newChannel(bytes))) {
VarCharVector vector = (VarCharVector) root.getVector(0);
vector.allocateNew(1);
- vector.setSafe(0,
"x".getBytes(java.nio.charset.StandardCharsets.UTF_8));
+ vector.setSafe(0, "x".getBytes(StandardCharsets.UTF_8));
root.setRowCount(1);
writer.start();
writer.writeBatch();
@@ -206,4 +330,12 @@ class AdbcObjectsReaderTest {
}
return new ArrowStreamReader(new
ByteArrayInputStream(bytes.toByteArray()), allocator);
}
+
+ private static byte[] utf8(String value) {
+ return value.getBytes(StandardCharsets.UTF_8);
+ }
+
+ private static List<Object[]> asList(Object value) {
+ return value == null ? null : (List<Object[]>) value;
+ }
}
diff --git
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcQueryBuilderNativeTest.java
b/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcQueryBuilderNativeTest.java
deleted file mode 100644
index 6c4f135b05a..00000000000
---
a/fe/fe-connector/fe-connector-adbc/src/test/java/org/apache/doris/connector/adbc/AdbcQueryBuilderNativeTest.java
+++ /dev/null
@@ -1,212 +0,0 @@
-// Licensed to the Apache Software Foundation (ASF) under one
-// or more contributor license agreements. See the NOTICE file
-// distributed with this work for additional information
-// regarding copyright ownership. The ASF licenses this file
-// to you under the Apache License, Version 2.0 (the
-// "License"); you may not use this file except in compliance
-// with the License. You may obtain a copy of the License at
-//
-// http://www.apache.org/licenses/LICENSE-2.0
-//
-// Unless required by applicable law or agreed to in writing,
-// software distributed under the License is distributed on an
-// "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY
-// KIND, either express or implied. See the License for the
-// specific language governing permissions and limitations
-// under the License.
-
-package org.apache.doris.connector.adbc;
-
-import org.apache.doris.connector.spi.ConnectorType;
-import org.apache.doris.connector.spi.handle.ConnectorColumnHandle;
-import org.apache.doris.connector.spi.handle.NamedColumnHandle;
-import org.apache.doris.connector.spi.pushdown.ConnectorAnd;
-import org.apache.doris.connector.spi.pushdown.ConnectorColumnRef;
-import org.apache.doris.connector.spi.pushdown.ConnectorComparison;
-import org.apache.doris.connector.spi.pushdown.ConnectorExpression;
-import org.apache.doris.connector.spi.pushdown.ConnectorIn;
-import org.apache.doris.connector.spi.pushdown.ConnectorIsNull;
-import org.apache.doris.connector.spi.pushdown.ConnectorLiteral;
-
-import org.apache.arrow.adbc.core.AdbcStatement;
-import org.apache.arrow.vector.ipc.ArrowReader;
-import org.junit.jupiter.api.Assertions;
-import org.junit.jupiter.api.Test;
-import org.junit.jupiter.api.io.TempDir;
-
-import java.nio.file.Path;
-import java.util.ArrayList;
-import java.util.List;
-import java.util.Map;
-import java.util.Optional;
-
-/**
- * Runs the generated SQL through the real SQLite ADBC driver.
- *
- * <p>Asserting the text of a statement only proves it is the text that was
intended. Whether a source
- * ACCEPTS it -- the quoting, the literal spellings, the placement of {@code
LIMIT} -- is a different
- * question, and one the ANSI dialect answers on behalf of sources nobody has
tried. This is the cheapest
- * available source that answers it for real.
- *
- * <p>Skips loudly without the native libraries; a skipped run has verified
nothing about real SQL
- * acceptance, only about string building.
- */
-class AdbcQueryBuilderNativeTest {
-
- private static final AdbcDialect ANSI =
AdbcDialectRegistry.defaultDialect();
- private static final AdbcTableHandle T1 =
- new AdbcTableHandle(new AdbcNamespace("main", ""), "t1");
-
- private static AdbcClient sqliteClient(Path dbFile) {
- return new AdbcClient(AdbcNativeTestSupport.sqliteDriver(),
"libadbc_driver_sqlite.so",
- null, "file:" + dbFile, null, null, Map.of());
- }
-
- /**
- * Values chosen so every literal form the dialect renders appears in a
predicate below, and so the
- * quoting matters: a column named with a reserved word and a string
holding a quote.
- */
- private static void seed(AdbcClient client) {
- client.withConnection(connection -> {
- for (String sql : new String[] {
- "CREATE TABLE t1 (id INTEGER, \"select\" REAL, name TEXT)",
- "INSERT INTO t1 VALUES (1, 1.5, 'a')",
- "INSERT INTO t1 VALUES (2, 2.5, 'O''Brien')",
- "INSERT INTO t1 VALUES (3, 3.5, NULL)"}) {
- try (AdbcStatement statement = connection.createStatement()) {
- statement.setSqlQuery(sql);
- statement.executeUpdate();
- }
- }
- return null;
- });
- }
-
- private static List<ConnectorColumnHandle> columns(String... names) {
- List<ConnectorColumnHandle> handles = new ArrayList<>(names.length);
- for (String name : names) {
- handles.add(new NamedColumnHandle(name));
- }
- return handles;
- }
-
- private static ConnectorColumnRef col(String name, String type) {
- return new ConnectorColumnRef(name, ConnectorType.of(type));
- }
-
- /** Runs the generated statement and returns its column names and row
count. */
- private static Result run(AdbcClient client, List<ConnectorColumnHandle>
cols,
- ConnectorExpression filter, long limit) {
- String sql = AdbcQueryBuilder.build(ANSI, T1, cols,
Optional.ofNullable(filter), limit).getSql();
- return client.withConnection(connection -> {
- try (AdbcStatement statement = connection.createStatement()) {
- statement.setSqlQuery(sql);
- try (AdbcStatement.QueryResult queryResult =
statement.executeQuery()) {
- ArrowReader reader = queryResult.getReader();
- List<String> names = new ArrayList<>();
- reader.getVectorSchemaRoot().getSchema().getFields()
- .forEach(field -> names.add(field.getName()));
- int rows = 0;
- while (reader.loadNextBatch()) {
- rows += reader.getVectorSchemaRoot().getRowCount();
- }
- return new Result(sql, names, rows);
- }
- }
- });
- }
-
- @Test
- void theSourceAcceptsAProjectionOfQuotedIdentifiers(@TempDir Path tempDir)
{
- try (AdbcClient client = sqliteClient(tempDir.resolve("scan.db"))) {
- seed(client);
- // "select" is a reserved word; unquoted it is a syntax error,
which is what makes this a real
- // test of the quoting rather than of the string building.
- Result result = run(client, columns("id", "select"), null, -1);
-
- Assertions.assertEquals(List.of("id", "select"),
result.columnNames, result.sql);
- Assertions.assertEquals(3, result.rows, result.sql);
- }
- }
-
- @Test
- void theSourceReturnsOnlyTheRequestedColumns(@TempDir Path tempDir) {
- // BE rejects any column it did not ask for, so this is the property
the scan depends on -- not
- // merely that the query runs.
- try (AdbcClient client = sqliteClient(tempDir.resolve("proj.db"))) {
- seed(client);
- Assertions.assertEquals(List.of("id"), run(client, columns("id"),
null, -1).columnNames);
- }
- }
-
- @Test
- void theSourceAcceptsTheNumericAndStringLiteralsTheDialectWrites(@TempDir
Path tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("lit.db"))) {
- seed(client);
-
- Assertions.assertEquals(2, run(client, columns("id"),
- new ConnectorComparison(ConnectorComparison.Operator.GT,
- col("id", "INT"), ConnectorLiteral.ofLong(1)),
-1).rows);
- Assertions.assertEquals(1, run(client, columns("id"),
- new ConnectorComparison(ConnectorComparison.Operator.LT,
- col("select", "DOUBLE"),
ConnectorLiteral.ofDouble(2.0d)), -1).rows);
- // The escaped quote survives the round trip to the source, rather
than ending the literal.
- Assertions.assertEquals(1, run(client, columns("id"),
- new ConnectorComparison(ConnectorComparison.Operator.EQ,
- col("name", "STRING"),
ConnectorLiteral.ofString("O'Brien")), -1).rows);
- }
- }
-
- @Test
- void theSourceAcceptsNullTestsInListsAndConjunctions(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("pred.db"))) {
- seed(client);
-
- Assertions.assertEquals(1, run(client, columns("id"),
- new ConnectorIsNull(col("name", "STRING"), false),
-1).rows);
- Assertions.assertEquals(2, run(client, columns("id"),
- new ConnectorIn(col("id", "INT"),
- List.of(ConnectorLiteral.ofLong(1),
ConnectorLiteral.ofLong(3)), false),
- -1).rows);
- Assertions.assertEquals(1, run(client, columns("id"), new
ConnectorAnd(List.of(
- new ConnectorComparison(ConnectorComparison.Operator.GE,
- col("id", "INT"), ConnectorLiteral.ofLong(2)),
- new ConnectorIsNull(col("name", "STRING"), true))),
-1).rows);
- }
- }
-
- @Test
- void theSourceAcceptsTheLimitClauseWhereTheBuilderPutsIt(@TempDir Path
tempDir) {
- try (AdbcClient client = sqliteClient(tempDir.resolve("limit.db"))) {
- seed(client);
- Assertions.assertEquals(2, run(client, columns("id"), null,
2).rows);
- Assertions.assertEquals(1, run(client, columns("id"),
- new ConnectorComparison(ConnectorComparison.Operator.GT,
- col("id", "INT"), ConnectorLiteral.ofLong(1)),
1).rows);
- }
- }
-
- @Test
- void theSourceAcceptsTheCountOnlyProjection(@TempDir Path tempDir) {
- // What a pushed-down COUNT(*) sends: one narrow column per row, no
table values.
- try (AdbcClient client = sqliteClient(tempDir.resolve("count.db"))) {
- seed(client);
- Result result = run(client, columns(), null, -1);
- Assertions.assertEquals(3, result.rows, result.sql);
- Assertions.assertEquals(1, result.columnNames.size(), result.sql);
- }
- }
-
- private static final class Result {
-
- private final String sql;
- private final List<String> columnNames;
- private final int rows;
-
- Result(String sql, List<String> columnNames, int rows) {
- this.sql = sql;
- this.columnNames = columnNames;
- this.rows = rows;
- }
- }
-}
diff --git
a/regression-test/suites/external_table_p0/adbc/test_adbc_catalog_scan.groovy
b/regression-test/suites/external_table_p0/adbc/test_adbc_catalog_scan.groovy
index 8b70e9b0629..699656e2dbc 100644
---
a/regression-test/suites/external_table_p0/adbc/test_adbc_catalog_scan.groovy
+++
b/regression-test/suites/external_table_p0/adbc/test_adbc_catalog_scan.groovy
@@ -80,10 +80,22 @@ suite("test_adbc_catalog_scan", "p0,external") {
String catalogName = "test_adbc_catalog_scan_catalog"
String dbName = "test_adbc_catalog_scan_db"
+ // A second database in the SOURCE, so the listing has a namespace it must
not mix into the other one.
+ // Created up front, so nothing below depends on a database appearing
behind Doris's back.
+ String otherDbName = "${dbName}_other"
sql """DROP CATALOG IF EXISTS ${catalogName}"""
sql """DROP DATABASE IF EXISTS ${dbName} FORCE"""
+ sql """DROP DATABASE IF EXISTS ${otherDbName} FORCE"""
sql """CREATE DATABASE ${dbName}"""
+ sql """CREATE DATABASE ${otherDbName}"""
+ sql """
+ CREATE TABLE ${otherDbName}.t_other (
+ `id` int NOT NULL
+ ) DISTRIBUTED BY HASH(`id`) BUCKETS 1
+ PROPERTIES ("replication_num" = "1")
+ """
+ sql """INSERT INTO ${otherDbName}.t_other VALUES (7)"""
sql """
CREATE TABLE ${dbName}.t1 (
@@ -125,6 +137,15 @@ suite("test_adbc_catalog_scan", "p0,external") {
def tableNames = sql("""SHOW TABLES FROM
${catalogName}.${dbName}""").collect { it[0] } as Set
assertTrue(tableNames.contains("t1"), "t1 missing from ${tableNames}")
assertFalse(tableNames.contains("v1"), "the view v1 was surfaced as a
table: ${tableNames}")
+ // getObjects filters are advisory, so a driver may answer a narrower
request with rows of another
+ // database too, and a reader that took the rows at face value would offer
that table as one of THIS
+ // database. The sibling's own listing is the other half of the same rule:
exactly its table, and none
+ // of this one's.
+ assertFalse(tableNames.contains("t_other"),
+ "a table of another database was listed under ${dbName}:
${tableNames}")
+ def otherTableNames = sql("""SHOW TABLES FROM
${catalogName}.${otherDbName}""").collect { it[0] } as Set
+ assertEquals(["t_other"] as Set, otherTableNames,
+ "the sibling database did not list exactly its own table:
${otherTableNames}")
// ---- scan ----
@@ -219,4 +240,5 @@ suite("test_adbc_catalog_scan", "p0,external") {
sql """DROP CATALOG ${singleRangeCatalog}"""
sql """DROP CATALOG ${catalogName}"""
sql """DROP DATABASE ${dbName} FORCE"""
+ sql """DROP DATABASE ${otherDbName} FORCE"""
}
diff --git
a/regression-test/suites/external_table_p0/adbc/test_adbc_metadata_ops.groovy
b/regression-test/suites/external_table_p0/adbc/test_adbc_metadata_ops.groovy
index 0d374b9d0df..70a1caa4214 100644
---
a/regression-test/suites/external_table_p0/adbc/test_adbc_metadata_ops.groovy
+++
b/regression-test/suites/external_table_p0/adbc/test_adbc_metadata_ops.groovy
@@ -194,6 +194,15 @@ suite("test_adbc_metadata_ops", "p0,external") {
sql """DESC ${catalogName}.${sqliteDb}.meta_a"""
sqliteExec("ALTER TABLE meta_a ADD COLUMN added_by_refresh_table TEXT;"
+ " UPDATE meta_a SET added_by_refresh_table = 'x';")
+
+ // The DESC above paid for this schema once. Until a REFRESH arrives
the connector keeps serving
+ // that copy, which is the half that gives the assertion below its
meaning -- a connector that
+ // re-read on every statement would satisfy that one while remembering
nothing. The DATABASE and
+ // CATALOG levels below assert only the positive half: they exercise
the same rule at a coarser key.
+ def columnsBeforeRefresh = sql("DESC
${catalogName}.${sqliteDb}.meta_a").collect { it[0] } as Set
+ assertFalse(columnsBeforeRefresh.contains("added_by_refresh_table"),
+ "the source's new column was visible before any REFRESH:
${columnsBeforeRefresh}")
+
sql """REFRESH TABLE ${catalogName}.${sqliteDb}.meta_a"""
def afterTableRefresh = sql("DESC
${catalogName}.${sqliteDb}.meta_a").collect { it[0] } as Set
assertTrue(afterTableRefresh.contains("added_by_refresh_table"),
diff --git
a/regression-test/suites/external_table_p0/adbc/test_adbc_predicate_pushdown.groovy
b/regression-test/suites/external_table_p0/adbc/test_adbc_predicate_pushdown.groovy
index f78e872128f..a7c0b7b5c25 100644
---
a/regression-test/suites/external_table_p0/adbc/test_adbc_predicate_pushdown.groovy
+++
b/regression-test/suites/external_table_p0/adbc/test_adbc_predicate_pushdown.groovy
@@ -171,6 +171,14 @@ suite("test_adbc_predicate_pushdown", "p0,external") {
return q
}
+ // The table itself is addressed through the REMOTE coordinates the
handle carries -- the source's
+ // database and table name -- and not through the Doris catalog the
user wrote. The backend that runs
+ // this statement has no Doris catalog in scope, so a statement
carrying those names, or none at all,
+ // addresses nothing. Quoting is left free: only the two-level shape
is pinned here.
+ String qualified = statementFor("id = 1")
+ assertTrue((qualified =~
/FROM\s+.?${dbName}.?\s*\.\s*.?t_pred.?/).find(),
+ "the remote statement does not address the source's database
and table: ${qualified}")
+
// Correctness, entirely independent of the above: the source
evaluates the same predicate itself.
// If pushdown ever translates a predicate WRONGLY -- as opposed to
not at all -- this is what
// catches it, and no baseline can absorb the difference.
diff --git
a/regression-test/suites/external_table_p0/adbc/test_adbc_sqlite_catalog_scan.groovy
b/regression-test/suites/external_table_p0/adbc/test_adbc_sqlite_catalog_scan.groovy
index 3a491ae5e55..75220f4d753 100644
---
a/regression-test/suites/external_table_p0/adbc/test_adbc_sqlite_catalog_scan.groovy
+++
b/regression-test/suites/external_table_p0/adbc/test_adbc_sqlite_catalog_scan.groovy
@@ -103,6 +103,8 @@ suite("test_adbc_sqlite_catalog_scan", "p0,external") {
INSERT INTO t1 VALUES (4, 'O''Brien', 4.5);
CREATE TABLE t2 (a INTEGER);
INSERT INTO t2 VALUES (10);
+ CREATE TABLE t_blob (id INTEGER, payload BLOB);
+ INSERT INTO t_blob VALUES (1, x'00ff');
CREATE VIEW v1 AS SELECT * FROM t1;
"""
@@ -136,8 +138,14 @@ suite("test_adbc_sqlite_catalog_scan", "p0,external") {
// ---- metadata ----
def databases = sql """SHOW DATABASES FROM ${catalogName}"""
- assertTrue(databases.any { it[0] == dbName },
+ def reportedNames = databases.collect { it[0].toString() }
+ assertTrue(reportedNames.contains(dbName),
"SHOW DATABASES did not report the SQLite namespace as 'main':
${databases}")
+ // ...and no second one was invented from the level SQLite does not
have. It reports the absent
+ // db_schema as the EMPTY STRING rather than as null, so a reader that
treats only null as absent
+ // produces a database whose name is "". The membership check above
would not notice it; this does.
+ assertFalse(reportedNames.contains(""),
+ "the empty schema name was turned into a database:
${reportedNames}")
// Views are excluded: the connector asks the source for base tables
only, so v1 must not be
// here even though it exists in the fixture.
@@ -149,6 +157,14 @@ suite("test_adbc_sqlite_catalog_scan", "p0,external") {
qt_desc_t1 """DESC ${catalogName}.${dbName}.t1"""
+ // A source BLOB column: the driver reports it as Arrow binary, and
Doris has no binary column type
+ // here, so it maps to the string type -- the same one a TEXT column
maps to, which is why the
+ // assertion is on what DESC renders rather than on a name that would
merely repeat the mapping.
+ def blobColumns = sql("""DESC ${catalogName}.${dbName}.t_blob""")
+ .collectEntries { [(it[0].toString()): it[1].toString()] }
+ assertEquals("text", blobColumns["payload"],
+ "a BLOB column was not mapped to the string type:
${blobColumns}")
+
// ---- reading ----
def allRows = sql """SELECT id, name, score FROM
${catalogName}.${dbName}.t1 ORDER BY id"""
@@ -323,6 +339,15 @@ suite("test_adbc_sqlite_catalog_scan", "p0,external") {
// cannot re-derive on its own: it caches the schema too, and clears
its copy either way.
sqliteExec("ALTER TABLE t_created_later ADD COLUMN extra TEXT;"
+ " UPDATE t_created_later SET extra = 'x';")
+
+ // The schema was read once, above. Until a REFRESH arrives the
connector keeps serving that copy,
+ // and this is the half that gives the assertion below its meaning: a
connector that re-read on every
+ // statement would satisfy that one while remembering nothing.
+ def columnsBeforeRefresh = sql("""DESC
${catalogName}.${dbName}.t_created_later""")
+ .collect { it[0] } as Set
+ assertFalse(columnsBeforeRefresh.contains("extra"),
+ "the source's new column was visible before any REFRESH:
${columnsBeforeRefresh}")
+
sql """REFRESH CATALOG ${catalogName}"""
def refreshedColumns = sql("""DESC
${catalogName}.${dbName}.t_created_later""")
.collect { it[0] } as Set
diff --git a/run-fe-ut.sh b/run-fe-ut.sh
index bb7b992135b..6b40bb4cb31 100755
--- a/run-fe-ut.sh
+++ b/run-fe-ut.sh
@@ -195,6 +195,9 @@ EXTRA_FE_MODULES="${EXTRA_FE_MODULES:-}"
parse_extra_fe_modules "${EXTRA_FE_MODULES}"
FE_MODULES=("fe-common" "fe-core")
+# fe-connector-adbc is not upstream of fe-core, so -am cannot discover its
tests
+# unless the module is named explicitly.
+FE_MODULES+=("fe-connector/fe-connector-adbc")
# The BE Java plugin modules. Nothing else runs these tests: no
be-java-extensions module is
# upstream of fe-core, so -am never reaches one, build.sh builds the reactor
with -DskipTests, and
# no GitHub workflow mentions the directory at all. What is in there is the
evidence that the
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]