This is an automated email from the ASF dual-hosted git repository.

davidzollo pushed a commit to branch dev
in repository https://gitbox.apache.org/repos/asf/seatunnel.git


The following commit(s) were added to refs/heads/dev by this push:
     new faf2400cd1 [Docs][Connector-V2] Clarify Kafka avro format vs Confluent 
Schema Registry (#11706)
faf2400cd1 is described below

commit faf2400cd1bf5c17cb9e6a61ba3d9151ac7d5fc7
Author: Daniel <[email protected]>
AuthorDate: Fri Aug 14 15:20:38 2026 +0800

    [Docs][Connector-V2] Clarify Kafka avro format vs Confluent Schema Registry 
(#11706)
    
    Co-authored-by: danielnadean 
<[email protected]>
    Co-authored-by: Claude Sonnet 5 <[email protected]>
---
 docs/en/connectors/source/Kafka.md | 2 ++
 docs/zh/connectors/source/Kafka.md | 2 ++
 2 files changed, 4 insertions(+)

diff --git a/docs/en/connectors/source/Kafka.md 
b/docs/en/connectors/source/Kafka.md
index 0de703876b..bfe4e8fd79 100644
--- a/docs/en/connectors/source/Kafka.md
+++ b/docs/en/connectors/source/Kafka.md
@@ -711,6 +711,8 @@ Note: the `key` field in NATIVE format is base64-encoded 
bytes.
 
 Kafka Source supports: `json`, `text`, `canal_json`, `debezium_json`, 
`ogg_json`, `avro`, `protobuf`, and `NATIVE`. Use `NATIVE` when you need access 
to Kafka-level metadata (headers, key, partition, timestamp) as part of the 
record.
 
+`format = avro` expects raw Avro-encoded messages. Unlike `protobuf` (see 
[Protobuf with Schema Registry wire 
format](#protobuf-with-schema-registry-wire-format)), there is no 
`strip_schema_registry_header`-equivalent option for `avro`: if a topic was 
produced by a Confluent `KafkaAvroSerializer` and its messages carry the 
Confluent Schema Registry wire-format header (magic byte + schema id), `format 
= avro` does not strip that header before deserializing, so reading will fail 
or produce [...]
+
 ### How do I configure SASL/Kerberos authentication?
 
 Pass authentication settings via `kafka.*` properties in the connector 
configuration:
diff --git a/docs/zh/connectors/source/Kafka.md 
b/docs/zh/connectors/source/Kafka.md
index 65f8cc350e..bda1300186 100644
--- a/docs/zh/connectors/source/Kafka.md
+++ b/docs/zh/connectors/source/Kafka.md
@@ -704,6 +704,8 @@ transform {
 
 支持:`json`、`text`、`canal_json`、`debezium_json`、`ogg_json`、`avro`、`protobuf` 和 
`NATIVE`。当需要将 Kafka 元数据(headers、key、partition、timestamp)作为记录字段使用时,选择 `NATIVE` 
格式。
 
+`format = avro` 仅支持原始(未经 Schema Registry 封装)的 Avro 消息。与 `protobuf`(参见 
[Protobuf with Schema Registry wire 
format](#protobuf-with-schema-registry-wire-format))不同,`avro` 没有对应 
`strip_schema_registry_header` 的选项:如果 topic 是由 Confluent `KafkaAvroSerializer` 
写入、消息中带有 Confluent Schema Registry 线格式头部(magic byte + schema id),`format = 
avro` 不会在反序列化前去除该头部,因此会读取失败或得到损坏的数据。`avro_schema` 仅用于为普通(非 Schema Registry)Avro 
消息提供 writer schema。
+
 ### 如何配置 SASL/Kerberos 认证?
 
 通过 `kafka.*` 属性传入认证参数:

Reply via email to