goutamadwant opened a new issue, #12304:
URL: https://github.com/apache/seatunnel/issues/12304
### Search before asking
- [x] Searched existing issues and pull requests for the same fix. Related
PR #11493 adds JSON logical-type cases in this file but does not address these
accounting paths.
### What happened
`SeaTunnelRow` byte accounting has two related defects for supported nested
values:
1. An array of maps containing a null element deserializes successfully from
JSON, but size accounting throws `NullPointerException`. Zeta's source
collector calculates this estimate before forwarding the record, so valid input
can fail during collection.
2. Schema-aware accounting omits the contents of arrays of rows. If the
outer row also has a nonzero-size field, the incomplete result is cached and
subsequently returned by the untyped method as well. Byte metrics and
byte-rate-limit permits can therefore exclude nested content.
These are accounting defects, not requests for a new nested-data type or a
change to serialization.
### SeaTunnel Version
Reproduced on `3.0.0-SNAPSHOT`, dev commit
`35b2716cde7d4c91a24fc618a8d9cae90e213db3`. Released versions were not
separately tested.
### Java or Scala Version
Reproduction and fix validation were run on **both Java 8 (Oracle 1.8.0_172)
and Java 11 (Temurin 11.0.19)**.
### Zeta or Flink or Spark Version
Zeta from the same checkout. Flink and Spark runtime tests were not run.
### Minimal API reproduction
Add the following JUnit tests to
`seatunnel-api/src/test/java/org/apache/seatunnel/api/table/type/SeaTunnelRowTest.java`
(the class already imports JUnit Assertions, Test, Collections, and Map):
```java
@Test
void nullableMapArrayAccounting() {
SeaTunnelRowType type = new SeaTunnelRowType(
new String[] {"items"},
new SeaTunnelDataType<?>[] {
new ArrayType<>(Map[].class,
new MapType<>(BasicType.STRING_TYPE,
BasicType.INT_TYPE))
});
SeaTunnelRow row = new SeaTunnelRow(new Object[] {
new Map[] {null, Collections.singletonMap("a", 1)}
});
Assertions.assertEquals(5, row.getBytesSize(type));
}
@Test
void rowArrayAccounting() {
SeaTunnelRowType elementType = new SeaTunnelRowType(
new String[] {"value"},
new SeaTunnelDataType<?>[] {BasicType.STRING_TYPE});
SeaTunnelRowType type = new SeaTunnelRowType(
new String[] {"id", "items"},
new SeaTunnelDataType<?>[] {
BasicType.INT_TYPE, new ArrayType<>(SeaTunnelRow[].class,
elementType)
});
Object[] fields = new Object[] {
7, new SeaTunnelRow[] {null, new SeaTunnelRow(new Object[] {"abcd"})}
};
SeaTunnelRow row = new SeaTunnelRow(fields);
Assertions.assertEquals(8, new SeaTunnelRow(fields).getBytesSize());
Assertions.assertEquals(8, row.getBytesSize(type));
Assertions.assertEquals(8, row.getBytesSize());
}
```
### Running Command
With `JAVA_HOME` set to each JDK, run:
```sh
./mvnw -B -pl seatunnel-api -am -Dtest=SeaTunnelRowTest \
-Dsurefire.failIfNoSpecifiedTests=false test
```
### Error Exception / before behavior
- Nullable map array: `java.lang.NullPointerException` when map-array
accounting dereferences the null element.
- Row array: `expected: <8> but was: <4>`. The typed result includes the
integer but omits `"abcd"`; the cached result remains 4. A fresh untyped
calculation returns 8.
- Additional regression coverage for a nested row array produces `expected:
<5> but was: <0>` before the fix.
The three added regression tests were run against unchanged production code
on both JDKs: 13 tests, 2 failures, 1 error; the existing 10 tests passed.
### SeaTunnel Config / runtime coverage
The fix also includes an embedded Zeta batch regression using this
configuration:
```hocon
env {
parallelism = 1
job.mode = "BATCH"
read_limit.bytes_per_second = 1000
}
source {
FakeSource {
row.num = 3
split.num = 1
schema = {
fields {
id = "int"
maps = "array<map<string,int>>"
}
}
rows = [
{ kind = INSERT, fields = { id = 1, maps = [null, {a = 1}] } },
{ kind = INSERT, fields = { id = 2, maps = [null] } },
{ kind = INSERT, fields = { id = 3, maps = [] } }
]
}
}
sink {
Console {
log.print.data = false
}
}
```
Run the supplied embedded test with:
```sh
./mvnw -B -pl seatunnel-engine/seatunnel-engine-server -am \
-Dtest=NestedRowAccountingTest -Dsurefire.failIfNoSpecifiedTests=false \
-Dskip.ui=true verify
```
After the fix, this job finishes with 3 records and 17 estimated bytes at
both source and sink. This was executed on Java 8 and Java 11.
The generic configuration type parser cannot currently express `ARRAY<ROW>`.
That case is tested through the schema API, JSON deserialization, and
single-table/multi-table source collectors, rather than an unsupported HOCON
example. A source-traced route also exists through Gravitino `list<struct>`
schema discovery into LocalFile JSON; an external Gravitino service was not
used for validation.
### Proposed fix and after behavior
- Skip null map-array elements while keeping the existing calculation for
non-null map entries.
- Include `ROW` in the existing recursive, schema-aware array calculation.
- Preserve nulls and all row payloads; only the size estimate changes.
- Add API, JSON, collector, and embedded batch regression coverage,
including empty/all-null arrays, nested arrays, and typed/untyped call order.
- Clarify estimated byte-limit accounting in the English and Chinese
documentation.
The production change is four lines in the shared row type. The nullable map
example returns 5 estimated bytes instead of throwing, and the row-array
example consistently returns 8 instead of 4.
### Advantages
- Prevents collection-time failures on valid nullable nested data.
- Counts nested row contents in existing byte metrics and byte-based flow
control.
- Fixes the shared calculation without connector-specific workarounds or new
dependencies.
### Breaking changes / compatibility
No public API, serialized field/layout, configuration name/default, or
dependency changes are proposed. No data payload is changed.
Observable behavior does change: affected jobs can report higher byte
metrics and consume more byte-limit permits, so throughput may decrease under
the same configured byte limit. This corrects under-accounting; it is not a
performance improvement claim. Existing size-estimation conventions remain
unchanged and are not exact serialized or network byte counts.
Mutable-row cache invalidation, integer bounds, cyclic object handling, and
configuration-parser support for row arrays are outside this fix.
### Validation
- Baseline reproduction on both Java 8 and Java 11.
- API/JSON dependency reactor: 523 tests passed per JDK, with no failures,
errors, or skips.
- New JSON and single-table/multi-table collector regressions; adjacent
collector, metrics, and rate-limiter checks passed.
- Embedded local batch test passed on both JDKs. The initial Java 8 fixture
setup error was corrected and that test was rerun successfully.
- Repository-wide formatting and Java 11 `./mvnw -q -DskipTests verify`
passed.
- Full connector E2E suites and external Gravitino validation were not run.
### Are you willing to submit PR?
- [x] Yes, the fix and regression tests are prepared.
### Code of Conduct
- [x] I agree to follow this project's [Code of
Conduct](https://www.apache.org/foundation/policies/conduct).
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]