This is an automated email from the ASF dual-hosted git repository.
morningman pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/doris-website.git
The following commit(s) were added to refs/heads/master by this push:
new f9d67b25c1f [docs] Add Apache Doris 4.1.4.1 release info (#4179)
f9d67b25c1f is described below
commit f9d67b25c1f777d39a7d414725034226d5cf7dcb
Author: Mingyu Chen (Rayner) <[email protected]>
AuthorDate: Tue Sep 29 22:39:14 2026 +0800
[docs] Add Apache Doris 4.1.4.1 release info (#4179)
## Summary
Publish the Apache Doris 4.1.4.1 hotfix as **Latest** on `/download/`
and `/releases/`.
4.1.4.1 is 4.1.4 plus four fixes: the `4.1.4...4.1.4.1` tag compare in
apache/doris is #68401, #68237, #68233 and #68507, plus two version
bumps, with nothing removed. The 4.1.4 binaries were already taken off
`/download/` in #4155 while its release notes stayed published. So the
4.1.4.1 notes are written as a supplement to 4.1.4 and link back to
them.
- `src/constant/download.data.ts`: add `4.1.4.1` as the first child of
the 4.1 branch with x64, x64-noavx2 and arm64 rows, so `ACTIVE_HEADS`
offers it as Latest. 4.1.4 stays absent from `ALL_VERSIONS`. The arm64
row is back because 4.1.4.1 fixes the aarch64 BE startup crash (#68507)
and ships an arm64 build.
- Add synchronized English and Chinese release notes from
[apache/doris#68518](https://github.com/apache/doris/issues/68518). The
source opens with a GitHub link to the 4.1.4 notes; on the site that
becomes an intro paragraph saying 4.1.4.1 is a hotfix for 4.1.4 that
includes everything in it, with a link to the on-site 4.1.4 release
notes.
- Update both Doris Core indexes (`releasenotes/core.md` and its zh-CN
mirror): the Latest tip now leads with 4.1.4.1 and also links the 4.1.4
notes, and a `2026-09-29` chronological entry is added. The 4.1.4 entry
stays.
- Add `v4.1/release-4.1.4.1` as the first item of the v4.1 release
sidebar.
`ACTIVE_CORE_BRANCHES` is unchanged.
## Release metadata
| | |
|---|---|
| Public version | 4.1.4.1 |
| Source filename version | `4.1.4.1` (no RC suffix) |
| Source directory |
`https://dist.apache.org/repos/dist/release/doris/4.1/4.1.4.1/` |
| Release date | 2026-09-29 (GitHub release publish date) |
| Position | latest |
| Branch change | none |
## Validation
- All 12 artifacts return 200: 3 source files on dist.apache.org and 9
binary files on OSS. Each `.sha512` file names the matching tarball.
- `node --test
doc-tools/skills/add-release/scripts/validate-release.test.mjs`: 6/6
passed.
- The bundled validator rejects four-part versions (`version must use
x.y.z format`), and its `ALL_VERSIONS` label scan only matches `x.y.z`.
A local copy with just those two regexes widened to accept an optional
fourth segment passes 65/65 checks, including the full 12-URL artifact
matrix. The validator itself is not changed in this PR.
- `sidebarsReleases.json` JSON parse check passed.
- `git diff --check` is clean.
## Build
Not run, per request: no `yarn build`, `yarn typecheck` or `yarn start`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
---------
Co-authored-by: morningman <[email protected]>
Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
---
docs/lakehouse/catalogs/fluss-catalog.md | 105 ++++++++++++++-------
.../current/core.md | 4 +-
.../current/v4.1/release-4.1.4.1.md | 27 ++++++
.../current/lakehouse/catalogs/fluss-catalog.md | 105 ++++++++++++++-------
releasenotes/core.md | 4 +-
releasenotes/v4.1/release-4.1.4.1.md | 27 ++++++
sidebarsReleases.json | 1 +
src/constant/download.data.ts | 34 +++++++
8 files changed, 239 insertions(+), 68 deletions(-)
diff --git a/docs/lakehouse/catalogs/fluss-catalog.md
b/docs/lakehouse/catalogs/fluss-catalog.md
index e8123d42240..3c91ec08c87 100644
--- a/docs/lakehouse/catalogs/fluss-catalog.md
+++ b/docs/lakehouse/catalogs/fluss-catalog.md
@@ -165,6 +165,8 @@ CREATE CATALOG fluss_sasl PROPERTIES (
| --- | --- |
| 5.0 | 1.0 |
+Fluss clusters older than 1.0 are not supported: the client APIs that Doris
relies on first appeared in Fluss 1.0.
+
## Metadata Mapping
<!-- Knowledge Type: Behavior Rules -->
@@ -173,7 +175,7 @@ The Fluss metadata hierarchy is Database -> Table, which
maps one to one onto Do
For a table with lake tiering enabled, `SHOW TABLES` lists only the table
itself. `tbl$lake` and `tbl$log` are two ways of reading the same table and do
not appear as separate entries.
-`SHOW CREATE TABLE` lists the Fluss table properties. They show whether lake
tiering is enabled (`table.datalake.enabled`) and which lake format is used
(`table.datalake.format`).
+`SHOW CREATE TABLE` shows the columns and comments of a Fluss table but not
its table properties. To find out whether a table has lake tiering enabled
(`table.datalake.enabled`) and which lake format it uses
(`table.datalake.format`), check the table definition on the Fluss side, for
example in Flink SQL. In Doris, only tables with lake tiering enabled have
`tbl$lake` and `tbl$log`.
## Column Type Mapping
@@ -251,8 +253,9 @@ SELECT * FROM fluss.db.`tbl$log`; -- Only the data in the
Fluss log that has n
**`tbl$lake`**
-- The Paimon Connector reads the Paimon table written by the tiering service
directly. Besides the table's own columns, it exposes three system columns that
Fluss adds when writing to the lake: `__bucket`, `__offset`, and `__timestamp`.
-- If the tiering service has not committed any data to the lake yet, the query
fails with `nothing has been tiered`.
+- The Paimon Connector reads the Paimon table written by the tiering service
directly. For a table tiered by Fluss 1.0, `tbl$lake` has exactly the same
columns as `tbl`. The `__bucket`, `__offset`, and `__timestamp` system columns
that earlier Fluss releases added to lake tables no longer exist.
+- It reads the lake at the readable lake snapshot that Fluss records, or at
the latest state of the lake when there is none yet. Fluss creates the Paimon
table together with the Fluss table, so before the tiering service commits
anything, `tbl$lake` returns an empty result instead of an error.
+- If Doris cannot find the lake table with the current lake connection
settings, the query fails with `nothing has been tiered`. This usually means
that settings such as `fluss.lake.paimon.warehouse` do not point to the
warehouse the Fluss cluster writes to. See Lake Connection Settings.
- The time travel and incremental queries of the Paimon Catalog are not
supported.
**`tbl$log`**
@@ -265,17 +268,17 @@ SELECT * FROM fluss.db.`tbl$log`; -- Only the data in
the Fluss log that has n
### Union Read Modes
<!-- Knowledge Type: Configuration Parameters + Behavior Rules -->
-<!-- Use Case: Controlling whether Union Read happens / Finding out why auto
mode fell back to a Fluss-only read -->
+<!-- Use Case: Controlling whether Union Read happens / Finding out why auto
mode fell back -->
Whether a query on the tiered lake table itself (`SELECT * FROM tbl`) does a
Union Read is controlled by the catalog property `fluss.union_read.mode`:
| Value | Behavior |
| --- | --- |
-| `auto` (default) | Does a Union Read when a readable lake snapshot exists.
Falls back to reading only Fluss when the conditions are not met. |
+| `auto` (default) | Does a Union Read when a readable lake snapshot exists.
When the conditions are not met, it falls back to reading the data from Fluss
instead of failing. See the cases below. |
| `required` | Union Read is mandatory. Fails when the conditions are not met
instead of falling back. |
-| `disabled` | No Union Read. Reads only Fluss and ignores the data in the
lake. |
+| `disabled` | No Union Read. Reads only the data currently in Fluss and
ignores the lake, including partitions that exist only in the lake. |
-The session variable `fluss_union_read_mode` overrides the catalog setting for
a single statement:
+The session variable `fluss_union_read_mode` overrides the catalog setting for
the current session:
```sql
SET fluss_union_read_mode = 'required'; -- No fallback to a Fluss-only read;
fail instead
@@ -283,19 +286,29 @@ SET fluss_union_read_mode = 'disabled'; -- Read from
Fluss only
SET fluss_union_read_mode = ''; -- Default: follow the catalog setting
```
-The Union Read mode only decides which path the data is read through; it does
not change the result. As long as the log in Fluss is still complete, all three
modes return the same rows. `required` is useful for regression tests and
troubleshooting, because it fails when the read path is not what you expect
instead of quietly taking another path.
+To change the mode for a single statement only, use the `SET_VAR` hint:
+
+```sql
+SELECT /*+ SET_VAR(fluss_union_read_mode='required') */ * FROM fluss.db.tbl;
+```
+
+The Union Read mode decides which path the data is read through. As long as
Fluss still holds all of the table's data, that is, no log has expired under
the TTL and no partition has been dropped from Fluss while it still exists in
the lake, all three modes return the same rows. `required` is useful for
regression tests and troubleshooting, because it fails when the read path is
not what you expect instead of quietly taking another path.
-**Cases where `auto` falls back to a Fluss-only read**
+**Cases where `auto` falls back**
| Case | Description |
| --- | --- |
-| No readable lake snapshot | The tiering service has not committed anything
yet, so all data is still in the log. |
-| `key-type` | The primary-key columns of a primary-key table have types that
cannot be compared exactly between the lake side and the log side: `FLOAT`,
`DOUBLE`, `TIMESTAMP`/`TIMESTAMP_LTZ` with precision greater than 6, and
`TIME`. |
+| No readable lake snapshot | The tiering service has not committed anything
yet, so all data is still in Fluss. The table is read from Fluss alone. |
+| `key-type` | A primary-key column of a primary-key table, other than a
partition column, has a type whose values cannot be compared exactly between
the lake side and the log side: `FLOAT`, `DOUBLE`, `TIME`, `TIMESTAMP` with
precision greater than 6, and `TIMESTAMP_LTZ`. When
`enable.mapping.timestamp_tz` is `false` (the default), a `TIMESTAMP_LTZ`
column of any precision falls in this case, because two different instants can
have the same local time during a daylight saving time overla [...]
| `partition-type` | A partition column of a primary-key table is not
`STRING`. Lake splits are matched to Fluss partitions by the text form of the
partition value, and only `STRING` guarantees that both sides write it the same
way. |
| `tail-truncated` | Part of the log after the lake snapshot of a primary-key
table has already been deleted by the Fluss TTL, so the log tail cannot be read
in full. |
| `tail-too-large` | The log tail of a primary-key table has more rows than
`fluss.union_read.max_tail_rows` or `fluss.union_read.max_total_tail_rows`
allows. |
-A primary-key table loses no data when it falls back to a Fluss-only read.
Fluss keeps the full state of a primary-key table; the query cannot use the
columnar files in the lake and runs slower. Log tables are different. A
Fluss-only read returns only what is still kept in the Fluss log, and once the
early log expires under the TTL, the result has fewer rows than a Union Read
would. `disabled` is not recommended for log tables that have been tiered.
+The last four cases apply to primary-key tables only, and `EXPLAIN` reports
them in the `degraded` field. When a primary-key table falls back, the
partitions that still exist in Fluss are read from Fluss in full. This loses no
data, because Fluss keeps the complete state of those partitions; the query
just cannot use the columnar files in the lake for them and runs slower.
Partitions that have been dropped from Fluss but still exist in the lake, for
example because of partition expiratio [...]
+
+For `partition-type`, Doris cannot reliably tell which Fluss partition a lake
split belongs to. If the lake holds a partition that Fluss no longer has, or a
partition value that the two sides write differently, the query fails with
`cannot be matched safely to a live fluss partition` instead of falling back.
Set the Union Read mode to `disabled` to read only the data in Fluss.
+
+Log tables are different. A Fluss-only read returns only what is still kept in
the Fluss log, and once the early log expires under the TTL, the result has
fewer rows than a Union Read would. `disabled` is not recommended for log
tables that have been tiered. A log table also has no fallback when its log
tail is incomplete: if part of the log after the lake snapshot has already
expired, a Union Read fails with `Part of the log tail has expired` in both
`auto` and `required` mode, and so d [...]
**Row limits for the log tail**
@@ -312,14 +325,14 @@ flussScan: readMode=default, unionRead=yes, lakeSplits=3,
suppressedLakeSplits=1
| Field | Meaning |
| --- | --- |
| `readMode` | `default` reads the whole table; `log` reads `tbl$log`. |
-| `unionRead` | Whether a Union Read was done. |
+| `unionRead` | Whether a Union Read was done. It is also `yes` for `tbl$log`,
whose starting offsets come from the lake snapshot; `lakeSplits` is 0 there. |
| `lakeSplits` | Number of lake splits. |
| `suppressedLakeSplits` | Number of those lake splits that have to be
filtered by the log tail. Primary-key tables only. |
| `logRanges` | Number of log scan ranges of a log table. |
| `pkRanges` | Number of buckets of a primary-key table read entirely from
Fluss (KV snapshot + change log). |
| `pkTailRanges` | Number of log tail scan ranges in a Union Read on a
primary-key table. |
| `mode` | The Union Read mode in effect. Carries a `(session)` suffix when it
comes from the session variable. |
-| `degraded` | Appears only when `auto` mode has fallen back to a Fluss-only
read. The value is the reason from the table above. |
+| `degraded` | Appears only when a primary-key table falls back in `auto`
mode. The value is one of `key-type`, `partition-type`, `tail-truncated`, and
`tail-too-large` from the table above. Having no readable lake snapshot is not
reported here; it only shows as `unionRead=no`. When partitions that exist only
in the lake are still read from the lake, `unionRead=yes` and `degraded` appear
together. |
### Lake Connection Settings
@@ -350,15 +363,16 @@ After the prefix, use Paimon's own parameter names, the
same as what follows `da
Overrides are applied per key, not as a whole. A key set in the catalog
overrides the same key reported by the Fluss cluster, and keys that are not set
still come from the cluster. If you only supply credentials, `warehouse`,
`metastore`, and the rest still come from the cluster.
-Storage parameters (credentials, endpoint, region, and addressing style) are
not handed to the Paimon Catalog. They are converted to Doris's own storage
parameter names and applied by the Doris storage layer. The reason is that
these settings serve both the FE (reading manifests) and the BE (reading data
files), and only the storage layer gives both sides the same set. The mapping
is as follows:
+Storage parameters (credentials, session tokens, endpoint, region, addressing
style, and Hadoop settings) are not handed to the Paimon Catalog. They are
converted to Doris's own storage parameter names and applied by the Doris
storage layer. The reason is that these settings serve both the FE (reading
manifests) and the BE (reading data files), and only the storage layer gives
both sides the same set. The mapping is as follows:
| Paimon parameter | Doris storage parameter |
| --- | --- |
-| `s3.access-key` / `s3.access.key` | `s3.access_key` |
-| `s3.secret-key` / `s3.secret.key` | `s3.secret_key` |
-| `s3.endpoint` | `s3.endpoint` |
-| `s3.region` | `s3.region` |
-| `s3.path-style-access` / `s3.path.style.access` | `use_path_style` |
+| `s3.access-key` / `s3.access.key` / `fs.s3a.access.key` | `s3.access_key` |
+| `s3.secret-key` / `s3.secret.key` / `fs.s3a.secret.key` | `s3.secret_key` |
+| `s3.session-token` / `s3.session.token` / `fs.s3a.session.token` |
`s3.session_token` |
+| `s3.endpoint` / `fs.s3a.endpoint` | `s3.endpoint` |
+| `s3.region` / `fs.s3a.endpoint.region` | `s3.region` |
+| `s3.path-style-access` / `s3.path.style.access` / `fs.s3a.path.style.access`
| `use_path_style` |
| `fs.oss.accessKeyId` | `oss.access_key` |
| `fs.oss.accessKeySecret` | `oss.secret_key` |
| `fs.oss.endpoint` | `oss.endpoint` |
@@ -366,6 +380,8 @@ Storage parameters (credentials, endpoint, region, and
addressing style) are not
| `fs.obs.secret.key` / `fs.obs.secret-key` | `obs.secret_key` |
| `fs.obs.endpoint` | `obs.endpoint` |
+Other parameters that start with `fs.`, `dfs.`, or `hadoop.`, such as HDFS HA
and Kerberos settings, are handed to the Doris storage layer as well, under
their original names.
+
`fluss.lake.paimon.metastore` accepts `filesystem` (default), `hive`, and
`rest`. Other parameters, such as the Hive Metastore address or the REST
authentication settings, are written the same way as in the [Paimon
Catalog](./paimon-catalog.mdx).
`ALTER CATALOG` can only set properties; it cannot remove them. To withdraw a
`fluss.lake.paimon.*` override, set it to an empty string, and Doris goes back
to the configuration reported by the Fluss cluster:
@@ -374,7 +390,7 @@ Storage parameters (credentials, endpoint, region, and
addressing style) are not
ALTER CATALOG fluss SET PROPERTIES ('fluss.lake.paimon.warehouse' = '');
```
-One Fluss Catalog reads lake tables with a single set of lake settings. After
the Fluss cluster changes `datalake.paimon.*`, run `REFRESH CATALOG` so that
Doris reloads it.
+One Fluss Catalog reads lake tables with a single set of lake settings. After
the Fluss cluster changes `datalake.paimon.*`, drop and recreate the catalog so
that Doris loads the new settings. `REFRESH CATALOG` only clears the metadata
cache and does not reload the lake settings.
### Query Profile
@@ -391,13 +407,22 @@ TableReader
├── FlussLakeReadTime 8s25ms Total time on the Paimon side,
including tail reads and filtering
│ ├── FlussLakeRangeNum 1
│ ├── FlussLakeSuppressRangeNum 2
-│ └── FlussLakeRowsReturned 7 Rows after filtering; add
SuppressedRows to get the count before filtering
+│ └── FlussLakeRowsReturned 7 Rows after filtering; add
FlussUnionSuppressedRows to get the count before filtering
├── FlussUnionTailReadTime 5s20ms One round trip to Fluss per bucket
└── FlussUnionSuppressTime 16.1us Filtering per block, grows with
the number of lake rows
```
Subtracting the time of each range type from the total of its side gives the
cost of initializing that reader (JNI class loading, or setting up the Paimon
read stack). A query that only reads a log table has no lake or primary-key
counters in its profile.
+A scan that reads lake splits also has the following counters under
`TableReader`. They carry values in a Union Read on a primary-key table:
+
+| Counter | Meaning |
+| --- | --- |
+| `FlussUnionSuppressedRows` | Lake rows dropped because the log tail updated
or deleted them. |
+| `FlussUnionTailKeysRead` | Change log records read from the log tails. |
+| `FlussUnionTailKeysRetained` | Distinct primary keys kept in memory to
filter the lake rows. |
+| `FlussUnionTailCacheHitCount` | Lake splits that reused a log tail already
read for the same bucket. |
+
## Query Performance
<!-- Knowledge Type: Architecture Principles + Performance Tuning -->
@@ -415,8 +440,10 @@ Which path is used depends on the type of the scan range,
which corresponds to t
| Log of a log table | `logRanges` | JNI |
| Whole primary-key table read (KV snapshot + change log) | `pkRanges` | JNI |
| Log tail of a primary-key table in a Union Read | `pkTailRanges` | JNI |
-| Lake splits of a log table, in a Union Read or `tbl$lake` | `lakeSplits` |
C++ native reader |
-| Lake splits of a primary-key table, in a Union Read or `tbl$lake` |
`lakeSplits`, `suppressedLakeSplits` | Decided by the Paimon Catalog rules, see
below. Filtering lake rows by the primary keys in the log tail happens in C++
on the BE |
+| Lake splits of a log table in a Union Read | `lakeSplits` | C++ native
reader |
+| Lake splits of a primary-key table in a Union Read | `lakeSplits`,
`suppressedLakeSplits` | Decided by the Paimon Catalog rules, see below.
Filtering lake rows by the primary keys in the log tail happens in C++ on the
BE |
+
+`tbl$lake` is planned by the Paimon Connector directly, so its `EXPLAIN` has
no `flussScan` line. It shows `paimonNativeReadSplits=N/M` instead, meaning N
of the M splits use the native reader, and its profile has no Fluss counters.
Lake splits are planned by the Paimon Connector, and whether they use the
native reader follows the same rules as a plain Paimon table:
@@ -426,9 +453,9 @@ Lake splits are planned by the Paimon Connector, and
whether they use the native
For a tiered table, Union Read lets the part in the lake (usually the vast
majority of the data) go through the native reader, and only the small amount
of log after the tiering offset goes through JNI. The more timely the tiering
(the smaller `table.datalake.freshness`), the less goes through JNI. That is
why tiered tables use Union Read by default. A few suggestions:
-- When analyzing history and freshness does not matter much, query `tbl$lake`
directly. The whole path is then the native reader.
+- When analyzing history and freshness does not matter much, query `tbl$lake`
directly. It skips the Fluss log entirely: a log table is then read wholly by
the native reader, and a primary-key table follows the Paimon rules above.
- `fluss.union_read.mode = disabled` sends the whole table through JNI. It is
not recommended except for troubleshooting.
-- When `auto` mode falls back to a Fluss-only read (`degraded=` appears in
`EXPLAIN`), the whole table also goes through JNI. For primary-key tables,
watch for the `key-type` and `partition-type` cases: they are determined by the
table schema, and only changing the table definition fixes them.
+- When `auto` mode falls back (`degraded=` appears in `EXPLAIN`), everything
read from Fluss also goes through JNI; only the partitions that exist only in
the lake are still read from the lake. Watch for the `key-type` and
`partition-type` cases: they are determined by the table schema, and only
changing the table definition fixes them.
- A Union Read on a primary-key table reads the primary keys in the log tail
once more for every bucket (`FlussUnionTailReadTime` in the profile). The
longer the tail, the higher the cost.
To confirm which path a query took, compare `lakeSplits` with `logRanges`,
`pkRanges`, and `pkTailRanges` in `EXPLAIN`, and compare `FlussLakeReadTime`
with `FlussLogReadTime` in the profile.
@@ -438,16 +465,16 @@ To confirm which path a query took, compare `lakeSplits`
with `logRanges`, `pkRa
<!-- Knowledge Type: Limitations -->
- Read only. `INSERT`, `UPDATE`, and `DELETE` are not supported, nor are
database and table management statements such as `CREATE TABLE` and `DROP
TABLE`.
-- Paimon is the only supported lake format. For tables whose
`table.datalake.format` is something else, the lake data cannot be read.
+- Paimon is the only supported lake format. For a table whose
`table.datalake.format` is something else, `tbl$lake` fails. Once such a table
has a readable lake snapshot, a query on the table itself also fails in `auto`
and `required` mode; set the Union Read mode to `disabled` to read the data in
Fluss.
- The Fluss `TIME` type maps to `UNSUPPORTED`, and that column cannot be
queried.
-- Predicates are not pushed down to Fluss or to the Paimon lake. Apart from
partition pruning, Doris applies all filters after reading the data. Column
pruning and sub-column pruning of nested types are both supported.
+- Predicates are not pushed down to Fluss: apart from partition pruning, Doris
filters the data read from Fluss after reading it. For the lake, the filter is
passed to Paimon when splits are planned, and Paimon uses it to skip partitions
and data files. Column pruning and sub-column pruning of nested types are both
supported.
- Time travel and incremental queries are not supported, including on
`tbl$lake`.
- Fluss provides only table-level row counts, without column-level statistics.
## FAQ
<!-- Knowledge Type: Troubleshooting -->
-<!-- Use Case: Missing Paimon plugin / Object storage credentials / Lake
snapshot not ready / Lake configuration changed / Unsupported partition column
type -->
+<!-- Use Case: Missing Paimon plugin / Object storage credentials / Lake
snapshot not ready / Lake table not found / Lake configuration changed /
Expired log tail / Unsupported partition column type / Partition matching
failure -->
1. Querying a tiered lake table fails with `the paimon connector plugin is not
available`
@@ -459,16 +486,28 @@ To confirm which path a query took, compare `lakeSplits`
with `logRanges`, `pkRa
3. Error `has no readable lake snapshot yet`
- The tiering service has not committed any data to the lake yet. In
`required` mode this is an error. Wait for the tiering service to commit, or
switch to `auto`.
+ The tiering service has not committed any data to the lake yet. A query on
the table itself reports this only in `required` mode: wait for the tiering
service to commit, or switch to `auto`. A query on `tbl$log` reports it in
every mode, because `tbl$log` starts from the lake snapshot; query the table
itself instead.
+
+4. `tbl$lake` fails with `nothing has been tiered`
+
+ Doris cannot find the lake table with the current lake connection
settings. Fluss 1.0 creates the lake table together with the Fluss table, so
this usually means that settings such as `fluss.lake.paimon.warehouse` do not
point to the warehouse the Fluss cluster writes to. See Lake Connection
Settings.
-4. Error `is already serving lake tables with a different paimon configuration`
+5. Error `is already serving lake tables with a different paimon configuration`
- The `datalake.paimon.*` configuration of the Fluss cluster changed after
the catalog was created. Run `REFRESH CATALOG` to reload it.
+ The `datalake.paimon.*` configuration of the Fluss cluster changed after
the catalog started reading lake tables. Drop and recreate the catalog.
`REFRESH CATALOG` does not reload the lake settings.
-5. Error `cannot be read: its partition column 'xx' has fluss type
TIMESTAMP(3)`
+6. Error `Part of the log tail has expired`
+
+ A Union Read on a log table, or a query on `tbl$log`, needs the log after
the lake snapshot, and part of it has already been deleted by the Fluss TTL.
Wait for Fluss to publish a newer readable lake snapshot. See Union Read Modes.
+
+7. Error `cannot be read: its partition column 'xx' has fluss type
TIMESTAMP(3)`
Doris cannot read partition columns of that type. See Partitioned Tables.
+8. Error `cannot be matched safely to a live fluss partition`
+
+ The primary-key table is partitioned by a column that is not `STRING`, and
the lake holds a partition that Doris cannot match to a partition in Fluss, for
example one that Fluss has already dropped. See the `partition-type` case in
Union Read Modes. To read the data still in Fluss, set the Union Read mode to
`disabled`.
+
## Debugging
<!-- Knowledge Type: Debugging Environment -->
diff --git a/i18n/zh-CN/docusaurus-plugin-content-docs-releases/current/core.md
b/i18n/zh-CN/docusaurus-plugin-content-docs-releases/current/core.md
index 51b67863d80..3faafa552fb 100644
--- a/i18n/zh-CN/docusaurus-plugin-content-docs-releases/current/core.md
+++ b/i18n/zh-CN/docusaurus-plugin-content-docs-releases/current/core.md
@@ -11,7 +11,7 @@
## Doris Core 版本发布说明
:::tip 最新发布
-🎉 4.1.4 版本已于 2026 年 09 月 10 日正式发布,详情可查看[版本发布](./v4.1/release-4.1.4.md)。Apache
Doris 4.1.4 新增 Merge-on-Write 表 ANN 索引、多模态文件 Embedding、面向大规模集群的自适应全局 Runtime
Filter 下发、OceanBase CDC Streaming Job 以及内部通信 TLS
支持,并修复查询执行、存储、导入、存算分离、湖仓一体与安全认证方面的多项问题。
+🎉 4.1.4.1 版本已于 2026 年 09 月 29
日正式发布,详情可查看[版本发布](./v4.1/release-4.1.4.1.md)。Apache Doris 4.1.4.1 是基于 4.1.4 的
Hotfix 版本,修复了逻辑 `OR` 处理可空 Boolean 值时 BE 崩溃、基表新增分区时分区 MTMV 执行全量刷新、含 `VARIANT` 列的
WAL 读取失败,以及 aarch64 BE 在 64 KiB 内存页的 Linux 系统上启动崩溃的问题。4.1.4 带来的 Merge-on-Write
表 ANN 索引、多模态文件 Embedding、OceanBase CDC Streaming Job、内部通信 TLS 等新功能与问题修复,详情可查看
[4.1.4 版本发布](./v4.1/release-4.1.4.md)。
<br />
🎉 4.0.8 版本已于 2026 年 08 月 14 日正式发布,详情可查看[版本发布](./v4.0/release-4.0.8.md)。Apache
Doris 4.0.8 是 4.0 系列维护版本,聚焦存算分离部署、导入与事务稳定性、File Cache 行为、内部接口安全加固以及湖仓兼容性。建议所有
4.0.x 用户升级。
@@ -30,6 +30,8 @@
<br />
+- [2026-09-29, Apache Doris 4.1.4.1 版本发布](./v4.1/release-4.1.4.1.md)
+
- [2026-09-10, Apache Doris 4.1.4 版本发布](./v4.1/release-4.1.4.md)
- [2026-08-14, Apache Doris 4.0.8 版本发布](./v4.0/release-4.0.8.md)
diff --git
a/i18n/zh-CN/docusaurus-plugin-content-docs-releases/current/v4.1/release-4.1.4.1.md
b/i18n/zh-CN/docusaurus-plugin-content-docs-releases/current/v4.1/release-4.1.4.1.md
new file mode 100644
index 00000000000..1c760292d87
--- /dev/null
+++
b/i18n/zh-CN/docusaurus-plugin-content-docs-releases/current/v4.1/release-4.1.4.1.md
@@ -0,0 +1,27 @@
+---
+{
+ "title": "Release 4.1.4.1",
+ "language": "zh-CN",
+ "description": "Apache Doris 4.1.4.1 版本发布说明"
+}
+---
+
+Apache Doris 4.1.4.1 是基于 4.1.4 的 Hotfix 版本,包含 4.1.4 的全部内容以及以下问题修复。4.1.4
的新功能、改进与问题修复请查看 [Apache Doris 4.1.4 版本发布说明](./release-4.1.4.md)。
+
+# 问题修复
+
+## 查询与执行
+
+- 修复逻辑 `OR` 处理可空的 Boolean 值时,多分支 `CASE WHEN` 等表达式导致 BE 崩溃的问题 (#68401)。
+
+## 物化视图
+
+- 修复基表新增分区时,分区 MTMV 执行全量刷新的问题。修复后,刷新只处理新增的分区 (#68237)。
+
+## 导入与 Streaming Job
+
+- 修复读取包含 `VARIANT` 列的 WAL 时,误报不支持的文件格式错误而导致读取失败的问题 (#68233)。
+
+## 平台
+
+- 将内置的 libunwind 从 1.6.2 升级到 1.8.3,修复 aarch64 BE 在使用 64 KiB 内存页的 Linux
系统上启动时崩溃的问题 (#68507)。
diff --git
a/i18n/zh-CN/docusaurus-plugin-content-docs/current/lakehouse/catalogs/fluss-catalog.md
b/i18n/zh-CN/docusaurus-plugin-content-docs/current/lakehouse/catalogs/fluss-catalog.md
index 7350ac64d37..ff39649dc81 100644
---
a/i18n/zh-CN/docusaurus-plugin-content-docs/current/lakehouse/catalogs/fluss-catalog.md
+++
b/i18n/zh-CN/docusaurus-plugin-content-docs/current/lakehouse/catalogs/fluss-catalog.md
@@ -165,6 +165,8 @@ CREATE CATALOG fluss_sasl PROPERTIES (
| --- | --- |
| 5.0 | 1.0 |
+不支持 1.0 之前的 Fluss 集群:Doris 依赖的客户端接口从 Fluss 1.0 才开始提供。
+
## 元数据映射
<!-- 知识类型: 行为规则 -->
@@ -173,7 +175,7 @@ Fluss 的元数据层级是 Database -> Table,和 Doris 一一对应,没有
开启了湖仓分层的表,`SHOW TABLES` 只列出表本身。`tbl$lake` 和 `tbl$log` 是同一张表的两种读法,不会单独出现在列表里。
-`SHOW CREATE TABLE` 会列出 Fluss
表的属性,从中能看到表有没有开湖仓分层(`table.datalake.enabled`)和湖格式(`table.datalake.format`)。
+`SHOW CREATE TABLE` 会显示 Fluss
表的列和注释,但不包含表属性。要确认表有没有开湖仓分层(`table.datalake.enabled`)、用的是哪种湖格式(`table.datalake.format`),需要到
Fluss 一侧查看表定义,比如在 Flink SQL 里查看。在 Doris 里,只有开启了湖仓分层的表才有 `tbl$lake` 和 `tbl$log`。
## 列类型映射
@@ -251,8 +253,9 @@ SELECT * FROM fluss.db.`tbl$log`; -- 只读 Fluss 日志中尚未分层的数
**`tbl$lake`**
-- 由 Paimon Connector 直接读分层服务写出的 Paimon 表。除了表本身的列,还多出 Fluss
写湖时附加的三个系统列:`__bucket`、`__offset`、`__timestamp`。
-- 分层服务还没往湖里提交过数据时,查询报错 `nothing has been tiered`。
+- 由 Paimon Connector 直接读分层服务写出的 Paimon 表。由 Fluss 1.0 分层的表,`tbl$lake` 的列和 `tbl`
完全相同。早期 Fluss 版本给湖表附加的 `__bucket`、`__offset`、`__timestamp` 三个系统列已经没有了。
+- 按 Fluss 记录的可读湖快照读湖;还没有可读湖快照时,读湖表的最新状态。Fluss 创建表时就会建好对应的 Paimon
表,所以分层服务提交数据之前,查 `tbl$lake` 返回空结果,不会报错。
+- Doris 按当前的湖连接配置找不到湖表时,查询报错 `nothing has been tiered`。这通常说明
`fluss.lake.paimon.warehouse` 等配置没有指向 Fluss 集群实际写入的仓库。参见【湖表连接配置】。
- 不支持 Paimon Catalog 的时间旅行和增量查询。
**`tbl$log`**
@@ -265,17 +268,17 @@ SELECT * FROM fluss.db.`tbl$log`; -- 只读 Fluss 日志中尚未分层的数
### Union Read 模式
<!-- 知识类型: 配置参数 + 行为规则 -->
-<!-- 适用场景: 控制是否 Union Read / 排查 auto 模式退回只读 Fluss 的原因 -->
+<!-- 适用场景: 控制是否 Union Read / 排查 auto 模式退回的原因 -->
查询湖仓分层表本身(即 `SELECT * FROM tbl`)时是否做 Union Read,由 Catalog 属性
`fluss.union_read.mode` 决定:
| 取值 | 行为 |
| --- | --- |
-| `auto`(默认) | 有可读的湖快照时做 Union Read;条件不满足时退回为只读 Fluss。 |
+| `auto`(默认) | 有可读的湖快照时做 Union Read;条件不满足时退回为从 Fluss 读取,不报错。具体情形见下文。 |
| `required` | 必须 Union Read,条件不满足时报错,不退回。 |
-| `disabled` | 不做 Union Read,只读 Fluss,不看湖里的数据。 |
+| `disabled` | 不做 Union Read,只读 Fluss 里当前的数据,不看湖里的数据,只存在于湖里的分区也读不到。 |
-会话变量 `fluss_union_read_mode` 可以在单条语句范围内覆盖 Catalog 的设置:
+会话变量 `fluss_union_read_mode` 可以在当前会话内覆盖 Catalog 的设置:
```sql
SET fluss_union_read_mode = 'required'; -- 不允许退回为只读 Fluss,退回时直接报错
@@ -283,19 +286,29 @@ SET fluss_union_read_mode = 'disabled'; -- 只从 Fluss 读取
SET fluss_union_read_mode = ''; -- 默认值,跟随 Catalog 的设置
```
-Union Read 模式只决定数据从哪条路径读出来,不改变结果:只要 Fluss 里的日志还完整,三种模式返回的行一样。`required`
适合回归测试和排查问题,读取路径不符合预期时直接报错,而不是换条路径把结果读出来。
+只想对单条语句生效时,用 `SET_VAR` Hint:
+
+```sql
+SELECT /*+ SET_VAR(fluss_union_read_mode='required') */ * FROM fluss.db.tbl;
+```
+
+Union Read 模式决定数据从哪条路径读出来。只要 Fluss 里还保留着表的全部数据,也就是没有日志按 TTL 过期,也没有分区已经从 Fluss
删掉而湖里还留着,三种模式返回的行就是一样的。`required` 适合回归测试和排查问题,读取路径不符合预期时直接报错,而不是换条路径把结果读出来。
-**`auto` 模式下退回为只读 Fluss 的情形**
+**`auto` 模式下的退回情形**
| 情形 | 说明 |
| --- | --- |
-| 没有可读的湖快照 | 分层服务还没提交过数据,此时全部数据都在日志里。 |
-| `key-type` | 主键表的主键列类型在湖端和日志端无法精确比较:`FLOAT`、`DOUBLE`、精度大于 6 的
`TIMESTAMP`/`TIMESTAMP_LTZ`,以及 `TIME`。 |
+| 没有可读的湖快照 | 分层服务还没提交过数据,此时全部数据都在 Fluss 里,整张表只从 Fluss 读取。 |
+| `key-type` | 主键表的某个主键列(分区列除外)是湖端和日志端无法精确比较的类型:`FLOAT`、`DOUBLE`、`TIME`、精度大于 6
的 `TIMESTAMP`,以及 `TIMESTAMP_LTZ`。`enable.mapping.timestamp_tz` 为
`false`(默认)时,任意精度的 `TIMESTAMP_LTZ` 列都属于这种情况,因为在夏令时回拨的重叠时段里,两个不同的时刻会对应同一个本地时间;为
`true` 时,只有精度大于 6 的才属于这种情况。 |
| `partition-type` | 主键表的分区列不是 `STRING`。湖端 Split 要靠分区值的文本形式匹配到 Fluss 分区,只有
`STRING` 能保证两边写法一致。 |
| `tail-truncated` | 主键表湖快照之后的日志已经被 Fluss 按 TTL 删掉了一部分,日志尾部读不全。 |
| `tail-too-large` | 主键表的日志尾部行数超过了 `fluss.union_read.max_tail_rows` 或
`fluss.union_read.max_total_tail_rows`。 |
-主键表退回为只读 Fluss 不丢数据:Fluss 保存着主键表的完整状态,只是用不上湖里的列存文件,读起来慢一些。日志表不同,只读 Fluss 拿到的是
Fluss 日志里还留着的数据,早期日志按 TTL 过期后,结果会比 Union Read 少。已经分层的日志表不建议用 `disabled`。
+后四种情形只针对主键表,`EXPLAIN` 会在 `degraded` 字段里给出原因。主键表退回时,Fluss 里还在的分区从 Fluss
完整读取,不丢数据,因为 Fluss 保存着这些分区的完整状态,只是这部分用不上湖里的列存文件,读起来慢一些;已经从 Fluss
删掉、但湖里还留着的分区(比如因分区过期被删),仍然从湖里读。未分区的表整张从 Fluss 读取。
+
+`partition-type` 情形下,Doris 没法可靠地判断湖端 Split 属于哪个 Fluss 分区。如果湖里有 Fluss
已经没有的分区,或者某个分区值在两边的写法不同,查询会报错 `cannot be matched safely to a live fluss
partition`,不会退回。这时可以把 Union Read 模式设为 `disabled`,只读 Fluss 里的数据。
+
+日志表不同,只读 Fluss 拿到的是 Fluss 日志里还留着的数据,早期日志按 TTL 过期后,结果会比 Union Read
少。已经分层的日志表不建议用 `disabled`。日志尾部不完整时,日志表也不会退回:湖快照之后的日志如果已经过期了一部分,`auto` 和
`required` 模式下的 Union Read 都会报错 `Part of the log tail has expired`,查询 `tbl$log`
在任何模式下也会报这个错。可以等 Fluss 发布更新的可读湖快照。
**日志尾部行数上限**
@@ -312,14 +325,14 @@ flussScan: readMode=default, unionRead=yes, lakeSplits=3,
suppressedLakeSplits=1
| 字段 | 含义 |
| --- | --- |
| `readMode` | `default` 表示读取整张表,`log` 表示读取的是 `tbl$log`。 |
-| `unionRead` | 是否做了 Union Read。 |
+| `unionRead` | 是否做了 Union Read。查 `tbl$log` 时也是 `yes`,因为它的起始位点来自湖快照,此时
`lakeSplits` 为 0。 |
| `lakeSplits` | 湖端 Split 数量。 |
| `suppressedLakeSplits` | 其中需要按日志尾部过滤的湖端 Split 数量,仅主键表有。 |
| `logRanges` | 日志表的日志扫描范围数量。 |
| `pkRanges` | 主键表从 Fluss 整体读取(KV 快照 + 变更日志)的 Bucket 数量。 |
| `pkTailRanges` | 主键表 Union Read 时,日志尾部的扫描范围数量。 |
| `mode` | 生效的 Union Read 模式。来自会话变量时带 `(session)` 后缀。 |
-| `degraded` | 只在 `auto` 模式退回为只读 Fluss 时出现,值是上表里的退回原因。 |
+| `degraded` | 只在 `auto` 模式下主键表退回时出现,值是上表中的
`key-type`、`partition-type`、`tail-truncated`、`tail-too-large`
之一。没有可读湖快照的情形不在这里体现,只表现为 `unionRead=no`。只存在于湖里的分区仍从湖里读时,`unionRead=yes` 和
`degraded` 会同时出现。 |
### 湖表连接配置
@@ -350,15 +363,16 @@ CREATE CATALOG fluss PROPERTIES (
覆盖是按 Key 进行的,不是整体替换。Catalog 里指定的 Key 覆盖 Fluss
集群上报的同名配置,没指定的仍然用集群的。只填凭证时,`warehouse`、`metastore` 等仍然来自集群。
-其中存储类参数(凭证、Endpoint、Region、寻址方式)不会交给 Paimon Catalog,而是转成 Doris 自己的存储参数名,由
Doris 的存储层统一生效。原因是这些配置要同时给 FE(读 Manifest)和 BE(读数据文件)用,只有走存储层两边才是同一套。对应关系如下:
+其中存储类参数(凭证、会话令牌、Endpoint、Region、寻址方式,以及 Hadoop 相关配置)不会交给 Paimon Catalog,而是转成
Doris 自己的存储参数名,由 Doris 的存储层统一生效。原因是这些配置要同时给 FE(读 Manifest)和
BE(读数据文件)用,只有走存储层两边才是同一套。对应关系如下:
| Paimon 参数写法 | 对应的 Doris 存储参数 |
| --- | --- |
-| `s3.access-key` / `s3.access.key` | `s3.access_key` |
-| `s3.secret-key` / `s3.secret.key` | `s3.secret_key` |
-| `s3.endpoint` | `s3.endpoint` |
-| `s3.region` | `s3.region` |
-| `s3.path-style-access` / `s3.path.style.access` | `use_path_style` |
+| `s3.access-key` / `s3.access.key` / `fs.s3a.access.key` | `s3.access_key` |
+| `s3.secret-key` / `s3.secret.key` / `fs.s3a.secret.key` | `s3.secret_key` |
+| `s3.session-token` / `s3.session.token` / `fs.s3a.session.token` |
`s3.session_token` |
+| `s3.endpoint` / `fs.s3a.endpoint` | `s3.endpoint` |
+| `s3.region` / `fs.s3a.endpoint.region` | `s3.region` |
+| `s3.path-style-access` / `s3.path.style.access` / `fs.s3a.path.style.access`
| `use_path_style` |
| `fs.oss.accessKeyId` | `oss.access_key` |
| `fs.oss.accessKeySecret` | `oss.secret_key` |
| `fs.oss.endpoint` | `oss.endpoint` |
@@ -366,6 +380,8 @@ CREATE CATALOG fluss PROPERTIES (
| `fs.obs.secret.key` / `fs.obs.secret-key` | `obs.secret_key` |
| `fs.obs.endpoint` | `obs.endpoint` |
+其他以 `fs.`、`dfs.`、`hadoop.` 开头的参数,比如 HDFS HA、Kerberos 相关配置,也交给 Doris
的存储层,参数名保持原样。
+
`fluss.lake.paimon.metastore` 支持 `filesystem`(默认)、`hive` 和 `rest`。其余参数,比如 Hive
Metastore 地址、REST 认证信息,按 [Paimon Catalog](./paimon-catalog.mdx) 的方式填写。
`ALTER CATALOG` 只能设置属性,不能删除属性。要撤销某个 `fluss.lake.paimon.*` 覆盖,把它设成空字符串,Doris
会重新使用 Fluss 集群上报的配置:
@@ -374,7 +390,7 @@ CREATE CATALOG fluss PROPERTIES (
ALTER CATALOG fluss SET PROPERTIES ('fluss.lake.paimon.warehouse' = '');
```
-一个 Fluss Catalog 只按一套湖配置读湖表。Fluss 集群改了 `datalake.paimon.*` 之后,需要执行 `REFRESH
CATALOG` 让 Doris 重新加载。
+一个 Fluss Catalog 只按一套湖配置读湖表。Fluss 集群改了 `datalake.paimon.*` 之后,需要删除并重新创建
Catalog,Doris 才会用上新的配置。`REFRESH CATALOG` 只清理元数据缓存,不会重新加载湖配置。
### 查询 Profile
@@ -391,13 +407,22 @@ TableReader
├── FlussLakeReadTime 8s25ms Paimon 侧总耗时,包含尾部读取与过滤
│ ├── FlussLakeRangeNum 1
│ ├── FlussLakeSuppressRangeNum 2
-│ └── FlussLakeRowsReturned 7 过滤后的行数;加上 SuppressedRows 即为过滤前的行数
+│ └── FlussLakeRowsReturned 7 过滤后的行数;加上 FlussUnionSuppressedRows
即为过滤前的行数
├── FlussUnionTailReadTime 5s20ms 每个 Bucket 到 Fluss 的一次往返
└── FlussUnionSuppressTime 16.1us 按 Block 过滤,随湖端行数增长
```
某一侧的总耗时减去它下面各范围类型的耗时,就是初始化这个读取器的开销(JNI 类加载,或 Paimon 读取栈的初始化)。只读日志表的查询,Profile
里不会出现湖和主键相关的计数器。
+读了湖端 Split 的扫描,`TableReader` 下还有下面几个计数器,对主键表做 Union Read 时才有实际数值:
+
+| 计数器 | 含义 |
+| --- | --- |
+| `FlussUnionSuppressedRows` | 因为被日志尾部更新或删除而丢掉的湖端行数。 |
+| `FlussUnionTailKeysRead` | 从日志尾部读到的变更日志记录数。 |
+| `FlussUnionTailKeysRetained` | 为过滤湖端行而留在内存里的不重复主键数。 |
+| `FlussUnionTailCacheHitCount` | 直接复用同一 Bucket 已读过的日志尾部的湖端 Split 数。 |
+
## 查询性能
<!-- 知识类型: 架构原理 + 性能调优 -->
@@ -415,8 +440,10 @@ BE 读 Fluss 表有两条路径,开销差别很大:
| 日志表的日志 | `logRanges` | JNI |
| 主键表整体读取(KV 快照 + 变更日志) | `pkRanges` | JNI |
| Union Read 中主键表的日志尾部 | `pkTailRanges` | JNI |
-| Union Read 或 `tbl$lake` 中,日志表的湖端 Split | `lakeSplits` | C++ native reader |
-| Union Read 或 `tbl$lake` 中,主键表的湖端 Split | `lakeSplits`、`suppressedLakeSplits`
| 按 Paimon Catalog 的规则决定,见下文。按日志尾部主键过滤湖端行的动作在 BE 的 C++ 里完成 |
+| Union Read 中日志表的湖端 Split | `lakeSplits` | C++ native reader |
+| Union Read 中主键表的湖端 Split | `lakeSplits`、`suppressedLakeSplits` | 按 Paimon
Catalog 的规则决定,见下文。按日志尾部主键过滤湖端行的动作在 BE 的 C++ 里完成 |
+
+`tbl$lake` 直接由 Paimon Connector 规划,`EXPLAIN` 里没有 `flussScan` 行,显示的是
`paimonNativeReadSplits=N/M`,表示 M 个 Split 中有 N 个走 native reader;它的 Profile 里也没有
Fluss 的计数器。
湖端 Split 由 Paimon Connector 规划,走不走 native reader 和查普通 Paimon 表时一样:
@@ -426,9 +453,9 @@ BE 读 Fluss 表有两条路径,开销差别很大:
对已分层的表,Union Read 让湖里的那部分(通常是绝大部分数据)走 native reader,只有分层位点之后的少量日志走
JNI。分层越及时(`table.datalake.freshness` 越小),走 JNI 的部分越少,这也是分层表默认采用 Union Read
的原因。几点建议:
-- 分析历史数据、对新鲜度要求不高时,直接查 `tbl$lake`,整条路径都是 native reader。
+- 分析历史数据、对新鲜度要求不高时,直接查 `tbl$lake`,完全不读 Fluss 日志:日志表全程走 native reader,主键表按上面的
Paimon 规则决定。
- `fluss.union_read.mode = disabled` 会让整张表改走 JNI,除了排查问题不建议使用。
-- `auto` 模式退回为只读 Fluss 时(`EXPLAIN` 里出现 `degraded=`),整张表同样改走了 JNI。主键表要留意
`key-type` 和 `partition-type` 两种情况,它们由表结构决定,得调整建表才能解决。
+- `auto` 模式退回时(`EXPLAIN` 里出现 `degraded=`),从 Fluss 读取的部分同样走
JNI,只有只存在于湖里的分区仍从湖里读。要留意 `key-type` 和 `partition-type` 两种情况,它们由表结构决定,得调整建表才能解决。
- 主键表的 Union Read 要为每个 Bucket 额外读一次日志尾部的主键(Profile 里的
`FlussUnionTailReadTime`),尾部越长开销越大。
要确认一条查询实际走了哪条路径,看 `EXPLAIN` 里 `lakeSplits` 与
`logRanges`、`pkRanges`、`pkTailRanges` 的比例,以及 Profile 里 `FlussLakeReadTime` 与
`FlussLogReadTime` 的耗时。
@@ -438,16 +465,16 @@ BE 读 Fluss 表有两条路径,开销差别很大:
<!-- 知识类型: 使用限制 -->
- 只读。不支持 `INSERT`、`UPDATE`、`DELETE`,也不支持 `CREATE TABLE`、`DROP TABLE` 这类库表管理操作。
-- 湖格式只支持 Paimon。`table.datalake.format` 是其他格式的表,读不了湖端数据。
+- 湖格式只支持 Paimon。`table.datalake.format` 是其他格式的表,查 `tbl$lake`
会报错;这类表有了可读的湖快照之后,在 `auto` 和 `required` 模式下查表本身也会报错,需要把 Union Read 模式设为
`disabled`,只读 Fluss 里的数据。
- Fluss 的 `TIME` 类型映射为 `UNSUPPORTED`,这一列查不了。
-- 谓词不下推到 Fluss 和 Paimon 湖端,除分区裁剪外,过滤都由 Doris 读到数据后再做。列裁剪和嵌套类型的子列裁剪都支持。
+- 谓词不下推到 Fluss:除分区裁剪外,从 Fluss 读出的数据由 Doris 读到之后再过滤。湖端规划 Split 时会把过滤条件交给
Paimon,由 Paimon 据此跳过不需要的分区和数据文件。列裁剪和嵌套类型的子列裁剪都支持。
- 不支持时间旅行和增量查询,`tbl$lake` 也不支持。
- Fluss 只提供表级行数,没有列级统计信息。
## 常见问题
<!-- 知识类型: 故障排查 -->
-<!-- 适用场景: Paimon 插件缺失 / 对象存储凭证 / 湖快照未就绪 / 湖配置变更 / 分区列类型不支持 -->
+<!-- 适用场景: Paimon 插件缺失 / 对象存储凭证 / 湖快照未就绪 / 找不到湖表 / 湖配置变更 / 日志尾部过期 / 分区列类型不支持 /
分区匹配失败 -->
1. 查询湖仓分层表时报错 `the paimon connector plugin is not available`
@@ -459,16 +486,28 @@ BE 读 Fluss 表有两条路径,开销差别很大:
3. 报错 `has no readable lake snapshot yet`
- 分层服务还没往湖里提交过数据。`required` 模式下这是错误,可以等分层服务提交,或者改用 `auto`。
+ 分层服务还没往湖里提交过数据。查表本身时,只有 `required` 模式会报这个错,可以等分层服务提交,或者改用 `auto`。查
`tbl$log` 时任何模式都会报这个错,因为 `tbl$log` 要从湖快照开始读,这时改查表本身即可。
+
+4. 查询 `tbl$lake` 报错 `nothing has been tiered`
+
+ Doris 按当前的湖连接配置找不到湖表。Fluss 1.0 创建表时就会建好湖表,所以这通常说明
`fluss.lake.paimon.warehouse` 等配置没有指向 Fluss 集群实际写入的仓库。参见【湖表连接配置】。
-4. 报错 `is already serving lake tables with a different paimon configuration`
+5. 报错 `is already serving lake tables with a different paimon configuration`
- Catalog 创建之后,Fluss 集群的 `datalake.paimon.*` 配置变了。执行 `REFRESH CATALOG` 重新加载。
+ Catalog 开始读湖表之后,Fluss 集群的 `datalake.paimon.*` 配置变了。删除并重新创建 Catalog
即可,`REFRESH CATALOG` 不会重新加载湖配置。
-5. 报错 `cannot be read: its partition column 'xx' has fluss type TIMESTAMP(3)`
+6. 报错 `Part of the log tail has expired`
+
+ 对日志表做 Union Read 或查询 `tbl$log` 时,需要读湖快照之后的日志,而其中一部分已经被 Fluss 按 TTL 删掉了。可以等
Fluss 发布更新的可读湖快照。参见【Union Read 模式】。
+
+7. 报错 `cannot be read: its partition column 'xx' has fluss type TIMESTAMP(3)`
分区列的类型 Doris 读不了。参见【分区表】。
+8. 报错 `cannot be matched safely to a live fluss partition`
+
+ 主键表的分区列不是 `STRING`,而湖里有 Doris 无法对应到 Fluss 分区的分区,比如 Fluss 里已经删掉的分区。参见【Union
Read 模式】中的 `partition-type`。要读 Fluss 里现有的数据,可以把 Union Read 模式设为 `disabled`。
+
## 功能调试
<!-- 知识类型: 调试环境 -->
diff --git a/releasenotes/core.md b/releasenotes/core.md
index ea2ae4def87..06f76f6c58b 100644
--- a/releasenotes/core.md
+++ b/releasenotes/core.md
@@ -11,7 +11,7 @@ This document presents Apache Doris Core release notes in
reverse chronological
## Doris Core Release Notes
:::tip Latest Release
-🎉 Version 4.1.4 is released. Check out the 🔗[Release
Notes](./v4.1/release-4.1.4.md) here. Apache Doris 4.1.4 adds ANN indexes on
Merge-on-Write tables, multimodal file embedding, adaptive global runtime
filter publishing for large clusters, OceanBase CDC streaming jobs, and TLS for
internal communication. It also includes fixes across query execution, storage,
load, cloud-native deployments, lakehouse, and authentication.
+🎉 Version 4.1.4.1 is released. Check out the 🔗[Release
Notes](./v4.1/release-4.1.4.1.md) here. Apache Doris 4.1.4.1 is a hotfix
release for 4.1.4. It fixes a BE crash when logical `OR` processes nullable
Boolean values, full refreshes of partitioned MTMVs when a base table adds a
partition, WAL read failures with `VARIANT` columns, and a BE startup crash on
aarch64 Linux systems with 64 KiB memory pages. For what 4.1.4 brings,
including ANN indexes on Merge-on-Write tables, multimodal fi [...]
<br />
@@ -34,6 +34,8 @@ This document presents Apache Doris Core release notes in
reverse chronological
<br />
+- [2026-09-29, Apache Doris 4.1.4.1 is released](./v4.1/release-4.1.4.1.md)
+
- [2026-09-10, Apache Doris 4.1.4 is released](./v4.1/release-4.1.4.md)
- [2026-08-14, Apache Doris 4.0.8 is released](./v4.0/release-4.0.8.md)
diff --git a/releasenotes/v4.1/release-4.1.4.1.md
b/releasenotes/v4.1/release-4.1.4.1.md
new file mode 100644
index 00000000000..9f877635231
--- /dev/null
+++ b/releasenotes/v4.1/release-4.1.4.1.md
@@ -0,0 +1,27 @@
+---
+{
+ "title": "Release 4.1.4.1",
+ "language": "en",
+ "description": "Here's the Apache Doris 4.1.4.1 release notes:"
+}
+---
+
+Apache Doris 4.1.4.1 is a hotfix release for 4.1.4. It includes everything in
4.1.4 plus the fixes below. For the new features, improvements, and bug fixes
in 4.1.4, see the [Apache Doris 4.1.4 release notes](./release-4.1.4.md).
+
+# Bugfix
+
+## Query & Execution
+
+- Fix a BE crash in expressions such as multi-branch `CASE WHEN` when logical
`OR` processes nullable Boolean values. (#68401)
+
+## Materialized Views
+
+- Fix partitioned MTMVs performing a full refresh when a base table adds a
partition. The refresh now processes only the new partition. (#68237)
+
+## Load & Streaming
+
+- Fix WAL reads with `VARIANT` columns incorrectly failing with an unsupported
file-format error. (#68233)
+
+## Platform
+
+- Fix an aarch64 BE startup crash on Linux systems with 64 KiB memory pages by
upgrading the bundled libunwind from 1.6.2 to 1.8.3. (#68507)
diff --git a/sidebarsReleases.json b/sidebarsReleases.json
index c0b97aeef49..4c70087517c 100644
--- a/sidebarsReleases.json
+++ b/sidebarsReleases.json
@@ -13,6 +13,7 @@
"type": "category",
"label": "v4.1",
"items": [
+ "v4.1/release-4.1.4.1",
"v4.1/release-4.1.4",
"v4.1/release-4.1.3",
"v4.1/release-4.1.2",
diff --git a/src/constant/download.data.ts b/src/constant/download.data.ts
index 1b3b4c06c42..cd3a7dc7f06 100644
--- a/src/constant/download.data.ts
+++ b/src/constant/download.data.ts
@@ -90,6 +90,40 @@ export const ALL_VERSIONS: AllVersionOption[] = [
label: '4.1',
value: '4.1',
children: [
+ {
+ label: '4.1.4.1',
+ value: '4.1.4.1',
+ majorVersion: '4.1',
+ items: [
+ {
+ label: CPUEnum.X64,
+ value: CPUEnum.X64,
+ gz: `${ORIGIN}apache-doris-4.1.4.1-bin-x64.tar.gz`,
+ asc:
`${ORIGIN}apache-doris-4.1.4.1-bin-x64.tar.gz.asc`,
+ sha512:
`${ORIGIN}apache-doris-4.1.4.1-bin-x64.tar.gz.sha512`,
+ source:
'https://dist.apache.org/repos/dist/release/doris/4.1/4.1.4.1/',
+ version: '4.1.4.1',
+ },
+ {
+ label: CPUEnum.X64NoAvx2,
+ value: CPUEnum.X64NoAvx2,
+ gz:
`${ORIGIN}apache-doris-4.1.4.1-bin-x64-noavx2.tar.gz`,
+ asc:
`${ORIGIN}apache-doris-4.1.4.1-bin-x64-noavx2.tar.gz.asc`,
+ sha512:
`${ORIGIN}apache-doris-4.1.4.1-bin-x64-noavx2.tar.gz.sha512`,
+ source:
'https://dist.apache.org/repos/dist/release/doris/4.1/4.1.4.1/',
+ version: '4.1.4.1',
+ },
+ {
+ label: CPUEnum.ARM64,
+ value: CPUEnum.ARM64,
+ gz: `${ORIGIN}apache-doris-4.1.4.1-bin-arm64.tar.gz`,
+ asc:
`${ORIGIN}apache-doris-4.1.4.1-bin-arm64.tar.gz.asc`,
+ sha512:
`${ORIGIN}apache-doris-4.1.4.1-bin-arm64.tar.gz.sha512`,
+ source:
'https://dist.apache.org/repos/dist/release/doris/4.1/4.1.4.1/',
+ version: '4.1.4.1',
+ },
+ ],
+ },
{
label: '4.1.3',
value: '4.1.3',
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]