goutamadwant opened a new issue, #12304:
URL: https://github.com/apache/seatunnel/issues/12304

   ### Search before asking
   
   - [x] Searched existing issues and pull requests for the same fix. Related 
PR #11493 adds JSON logical-type cases in this file but does not address these 
accounting paths.
   
   ### What happened
   
   `SeaTunnelRow` byte accounting has two related defects for supported nested 
values:
   
   1. An array of maps containing a null element deserializes successfully from 
JSON, but size accounting throws `NullPointerException`. Zeta's source 
collector calculates this estimate before forwarding the record, so valid input 
can fail during collection.
   2. Schema-aware accounting omits the contents of arrays of rows. If the 
outer row also has a nonzero-size field, the incomplete result is cached and 
subsequently returned by the untyped method as well. Byte metrics and 
byte-rate-limit permits can therefore exclude nested content.
   
   These are accounting defects, not requests for a new nested-data type or a 
change to serialization.
   
   ### SeaTunnel Version
   
   Reproduced on `3.0.0-SNAPSHOT`, dev commit 
`35b2716cde7d4c91a24fc618a8d9cae90e213db3`. Released versions were not 
separately tested.
   
   ### Java or Scala Version
   
   Reproduction and fix validation were run on **both Java 8 (Oracle 1.8.0_172) 
and Java 11 (Temurin 11.0.19)**.
   
   ### Zeta or Flink or Spark Version
   
   Zeta from the same checkout. Flink and Spark runtime tests were not run.
   
   ### Minimal API reproduction
   
   Add the following JUnit tests to 
`seatunnel-api/src/test/java/org/apache/seatunnel/api/table/type/SeaTunnelRowTest.java`
 (the class already imports JUnit Assertions, Test, Collections, and Map):
   
   ```java
   @Test
   void nullableMapArrayAccounting() {
       SeaTunnelRowType type = new SeaTunnelRowType(
               new String[] {"items"},
               new SeaTunnelDataType<?>[] {
                   new ArrayType<>(Map[].class,
                           new MapType<>(BasicType.STRING_TYPE, 
BasicType.INT_TYPE))
               });
       SeaTunnelRow row = new SeaTunnelRow(new Object[] {
           new Map[] {null, Collections.singletonMap("a", 1)}
       });
       Assertions.assertEquals(5, row.getBytesSize(type));
   }
   
   @Test
   void rowArrayAccounting() {
       SeaTunnelRowType elementType = new SeaTunnelRowType(
               new String[] {"value"},
               new SeaTunnelDataType<?>[] {BasicType.STRING_TYPE});
       SeaTunnelRowType type = new SeaTunnelRowType(
               new String[] {"id", "items"},
               new SeaTunnelDataType<?>[] {
                   BasicType.INT_TYPE, new ArrayType<>(SeaTunnelRow[].class, 
elementType)
               });
       Object[] fields = new Object[] {
           7, new SeaTunnelRow[] {null, new SeaTunnelRow(new Object[] {"abcd"})}
       };
       SeaTunnelRow row = new SeaTunnelRow(fields);
       Assertions.assertEquals(8, new SeaTunnelRow(fields).getBytesSize());
       Assertions.assertEquals(8, row.getBytesSize(type));
       Assertions.assertEquals(8, row.getBytesSize());
   }
   ```
   
   ### Running Command
   
   With `JAVA_HOME` set to each JDK, run:
   
   ```sh
   ./mvnw -B -pl seatunnel-api -am -Dtest=SeaTunnelRowTest \
     -Dsurefire.failIfNoSpecifiedTests=false test
   ```
   
   ### Error Exception / before behavior
   
   - Nullable map array: `java.lang.NullPointerException` when map-array 
accounting dereferences the null element.
   - Row array: `expected: <8> but was: <4>`. The typed result includes the 
integer but omits `"abcd"`; the cached result remains 4. A fresh untyped 
calculation returns 8.
   - Additional regression coverage for a nested row array produces `expected: 
<5> but was: <0>` before the fix.
   
   The three added regression tests were run against unchanged production code 
on both JDKs: 13 tests, 2 failures, 1 error; the existing 10 tests passed.
   
   ### SeaTunnel Config / runtime coverage
   
   The fix also includes an embedded Zeta batch regression using this 
configuration:
   
   ```hocon
   env {
     parallelism = 1
     job.mode = "BATCH"
     read_limit.bytes_per_second = 1000
   }
   source {
     FakeSource {
       row.num = 3
       split.num = 1
       schema = {
         fields {
           id = "int"
           maps = "array<map<string,int>>"
         }
       }
       rows = [
         { kind = INSERT, fields = { id = 1, maps = [null, {a = 1}] } },
         { kind = INSERT, fields = { id = 2, maps = [null] } },
         { kind = INSERT, fields = { id = 3, maps = [] } }
       ]
     }
   }
   sink {
     Console {
       log.print.data = false
     }
   }
   ```
   
   Run the supplied embedded test with:
   
   ```sh
   ./mvnw -B -pl seatunnel-engine/seatunnel-engine-server -am \
     -Dtest=NestedRowAccountingTest -Dsurefire.failIfNoSpecifiedTests=false \
     -Dskip.ui=true verify
   ```
   
   After the fix, this job finishes with 3 records and 17 estimated bytes at 
both source and sink. This was executed on Java 8 and Java 11.
   
   The generic configuration type parser cannot currently express `ARRAY<ROW>`. 
That case is tested through the schema API, JSON deserialization, and 
single-table/multi-table source collectors, rather than an unsupported HOCON 
example. A source-traced route also exists through Gravitino `list<struct>` 
schema discovery into LocalFile JSON; an external Gravitino service was not 
used for validation.
   
   ### Proposed fix and after behavior
   
   - Skip null map-array elements while keeping the existing calculation for 
non-null map entries.
   - Include `ROW` in the existing recursive, schema-aware array calculation.
   - Preserve nulls and all row payloads; only the size estimate changes.
   - Add API, JSON, collector, and embedded batch regression coverage, 
including empty/all-null arrays, nested arrays, and typed/untyped call order.
   - Clarify estimated byte-limit accounting in the English and Chinese 
documentation.
   
   The production change is four lines in the shared row type. The nullable map 
example returns 5 estimated bytes instead of throwing, and the row-array 
example consistently returns 8 instead of 4.
   
   ### Advantages
   
   - Prevents collection-time failures on valid nullable nested data.
   - Counts nested row contents in existing byte metrics and byte-based flow 
control.
   - Fixes the shared calculation without connector-specific workarounds or new 
dependencies.
   
   ### Breaking changes / compatibility
   
   No public API, serialized field/layout, configuration name/default, or 
dependency changes are proposed. No data payload is changed.
   
   Observable behavior does change: affected jobs can report higher byte 
metrics and consume more byte-limit permits, so throughput may decrease under 
the same configured byte limit. This corrects under-accounting; it is not a 
performance improvement claim. Existing size-estimation conventions remain 
unchanged and are not exact serialized or network byte counts.
   
   Mutable-row cache invalidation, integer bounds, cyclic object handling, and 
configuration-parser support for row arrays are outside this fix.
   
   ### Validation
   
   - Baseline reproduction on both Java 8 and Java 11.
   - API/JSON dependency reactor: 523 tests passed per JDK, with no failures, 
errors, or skips.
   - New JSON and single-table/multi-table collector regressions; adjacent 
collector, metrics, and rate-limiter checks passed.
   - Embedded local batch test passed on both JDKs. The initial Java 8 fixture 
setup error was corrected and that test was rerun successfully.
   - Repository-wide formatting and Java 11 `./mvnw -q -DskipTests verify` 
passed.
   - Full connector E2E suites and external Gravitino validation were not run.
   
   ### Are you willing to submit PR?
   
   - [x] Yes, the fix and regression tests are prepared.
   
   ### Code of Conduct
   
   - [x] I agree to follow this project's [Code of 
Conduct](https://www.apache.org/foundation/policies/conduct).
   
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to