This is an automated email from the ASF dual-hosted git repository.

github-merge-queue[bot] pushed a commit to branch 
gh-readonly-queue/main/pr-24811-40488988ad596c9b093ad60e1453430d803ce33c
in repository https://gitbox.apache.org/repos/asf/datafusion.git

commit 690ec4c997ed0f87fdb59e92ddcd1d31e5db81b8
Author: Yongting You <[email protected]>
AuthorDate: Wed Sep 9 07:17:33 2026 +0000

    Doc: note `enforce_batch_size_in_joins` only applies to symmetric hash join 
(#24811)
    
    ## Which issue does this PR close?
    
    <!--
    We generally require a GitHub issue to be filed for all bug fixes and
    enhancements and this helps us generate change logs for our releases.
    You can link an issue to this PR using the GitHub syntax. For example
    `Closes #123` indicates that this PR will close issue #123.
    -->
    
    - Closes #.
    
    ## Rationale for this change
    
    <!--
    Why are you proposing this change? If this is already explained clearly
    in the issue then this section is not needed.
    Explaining clearly why changes are proposed helps reviewers understand
    your changes and offer better suggestions for fixes.
    
    Please explain the problem you are trying to solve in terms of the
    user-visible
    behavior, rather than the implementation.
    
    For example, "The code in `foo.rs` doesn't handle nulls" is a symptom of
    the
    implementation. "COUNT(DISTINCT) returns wrong results when the column
    contains
    nulls" is the user-visible problem.
    -->
    This config used to be applicable to most join operators
    - https://github.com/apache/datafusion/pull/12969
    
    Now only SHJ is using it, this PR updates the doc. (I think we can
    eventually deprecate it, but it requires further investigation, so it's
    left to future work, tracked in
    https://github.com/apache/datafusion/issues/24812)
    
    To verify, search for the config name in the codebase, and only SHJ is
    using it.
    
    ## What changes are included in this PR?
    
    <!--
    There is no need to duplicate the description in the issue here, but it
    is sometimes worth providing a summary of the individual changes in this
    PR.
    -->
    
    ## What is the testing strategy for this PR?
    
    <!--
    We typically require tests for all PRs in order to:
    1. Prevent the code from being accidentally broken by subsequent changes
    2. Serve as another way to document the expected behavior of the code
    
    Briefly describe how this PR is tested, and point to the specific tests
    you added. For example: 'This new feature is covered by the
    `sqllogictest` cases added in `foo.slt`'.
    
    If this PR does not add tests, explain why. For example, if the change
    is already covered by existing tests, please mention it.
    
    You should also check the `codecov` bot reply on this PR to confirm the
    changed code is exercised.
    -->
    
    ## Are there any user-facing changes?
    
    <!--
    If there are user-facing changes then we may require documentation to be
    updated before approving the PR.
    
    If there are any breaking changes to public APIs, please add the `api
    change` label.
    -->
---
 datafusion/common/src/config.rs                           | 1 +
 datafusion/sqllogictest/test_files/information_schema.slt | 2 +-
 docs/source/user-guide/configs.md                         | 2 +-
 3 files changed, 3 insertions(+), 2 deletions(-)

diff --git a/datafusion/common/src/config.rs b/datafusion/common/src/config.rs
index f770cebf88..cabddfc2c2 100644
--- a/datafusion/common/src/config.rs
+++ b/datafusion/common/src/config.rs
@@ -1135,6 +1135,7 @@ config_namespace! {
         /// DataFusion will not enforce batch size in joins. Enforcing batch 
size
         /// in joins can reduce memory usage when joining large
         /// tables with a highly-selective join filter, but is also slightly 
slower.
+        /// Note: this option currently only applies to the symmetric hash 
join.
         pub enforce_batch_size_in_joins: bool, default = false
 
         /// Size (bytes) of data buffer DataFusion uses when writing output 
files.
diff --git a/datafusion/sqllogictest/test_files/information_schema.slt 
b/datafusion/sqllogictest/test_files/information_schema.slt
index 4551f21f21..c5aaf02500 100644
--- a/datafusion/sqllogictest/test_files/information_schema.slt
+++ b/datafusion/sqllogictest/test_files/information_schema.slt
@@ -383,7 +383,7 @@ datafusion.execution.enable_file_stream_work_stealing true 
When `true` (the defa
 datafusion.execution.enable_migration_aggregate true Whether aggregation uses 
the implementation from the major refactor completed in the 56.0.0 release. 
When set to `false`, aggregation falls back to the implementation used before 
55.0.0. The fallback exists only as a workaround for bugs in the new 
implementation and will be removed, together with this option, after the 56.0.0 
release. See <https://github.com/apache/datafusion/issues/22710> for details.
 datafusion.execution.enable_nlj_coordinated_fallback true Enables the 
memory-limited fallback for `NestedLoopJoinExec` join types that emit unmatched 
left rows in the final output (LEFT, LEFT SEMI, LEFT ANTI, LEFT MARK, FULL) 
when the right side has multiple partitions. This fallback coordinates 
per-chunk left state (visited bitmap and probe-thread counter) across all 
right-side partitions, which assumes every partition runs in the same process. 
Distributed engines that execute each outp [...]
 datafusion.execution.enable_recursive_ctes true Should DataFusion support 
recursive CTEs
-datafusion.execution.enforce_batch_size_in_joins false Should DataFusion 
enforce batch size in joins or not. By default, DataFusion will not enforce 
batch size in joins. Enforcing batch size in joins can reduce memory usage when 
joining large tables with a highly-selective join filter, but is also slightly 
slower.
+datafusion.execution.enforce_batch_size_in_joins false Should DataFusion 
enforce batch size in joins or not. By default, DataFusion will not enforce 
batch size in joins. Enforcing batch size in joins can reduce memory usage when 
joining large tables with a highly-selective join filter, but is also slightly 
slower. Note: this option currently only applies to the symmetric hash join.
 datafusion.execution.hash_join_buffering_capacity 0 How many bytes to buffer 
in the probe side of hash joins while the build side is concurrently being 
built. Without this, hash joins will wait until the full materialization of the 
build side before polling the probe side. This is useful in scenarios where the 
query is not completely CPU bounded, allowing to do some early work 
concurrently and reducing the latency of the query. Note that when hash join 
buffering is enabled, the probe sid [...]
 datafusion.execution.keep_partition_by_columns false Should DataFusion keep 
the columns used for partition_by in the output RecordBatches
 datafusion.execution.listing_table_factory_infer_partitions true Should a 
`ListingTable` created through the `ListingTableFactory` infer table partitions 
from Hive compliant directories. Defaults to true (partition columns are 
inferred and will be represented in the table schema).
diff --git a/docs/source/user-guide/configs.md 
b/docs/source/user-guide/configs.md
index fbeb38f1ce..361f8882f2 100644
--- a/docs/source/user-guide/configs.md
+++ b/docs/source/user-guide/configs.md
@@ -141,7 +141,7 @@ The following configuration settings are available:
 | datafusion.execution.skip_partial_aggregation_probe_ratio_threshold     | 
0.8                       | Aggregation ratio (number of distinct groups / 
number of input rows) threshold for skipping partial aggregation. If the value 
is greater then partial aggregation will skip aggregation for further input     
                                                                                
                                                                                
                       [...]
 | datafusion.execution.skip_partial_aggregation_probe_rows_threshold      | 
100000                    | Number of input rows partial aggregation partition 
should process, before aggregation ratio check and trying to switch to skipping 
aggregation mode                                                                
                                                                                
                                                                                
                  [...]
 | datafusion.execution.use_row_number_estimates_to_optimize_partitioning  | 
false                     | Should DataFusion use row number estimates at the 
input to decide whether increasing parallelism is beneficial or not. By 
default, only exact row numbers (not estimates) are used for this decision. 
Setting this flag to `true` will likely produce better plans. if the source of 
statistics is accurate. We plan to make this the default in the future.         
                                [...]
-| datafusion.execution.enforce_batch_size_in_joins                        | 
false                     | Should DataFusion enforce batch size in joins or 
not. By default, DataFusion will not enforce batch size in joins. Enforcing 
batch size in joins can reduce memory usage when joining large tables with a 
highly-selective join filter, but is also slightly slower.                      
                                                                                
                           [...]
+| datafusion.execution.enforce_batch_size_in_joins                        | 
false                     | Should DataFusion enforce batch size in joins or 
not. By default, DataFusion will not enforce batch size in joins. Enforcing 
batch size in joins can reduce memory usage when joining large tables with a 
highly-selective join filter, but is also slightly slower. Note: this option 
currently only applies to the symmetric hash join.                              
                              [...]
 | datafusion.execution.objectstore_writer_buffer_size                     | 
10485760                  | Size (bytes) of data buffer DataFusion uses when 
writing output files. This affects the size of the data chunks that are 
uploaded to remote object stores (e.g. AWS S3). If very large (>= 100 GiB) 
output files are being written, it may be necessary to increase this size to 
avoid errors from the remote end point.                                         
                                    [...]
 | datafusion.execution.enable_ansi_mode                                   | 
false                     | Whether to enable ANSI SQL mode. The flag is 
experimental and relevant only for DataFusion Spark built-in functions When 
`enable_ansi_mode` is set to `true`, the query engine follows ANSI SQL 
semantics for expressions, casting, and error handling. This means: - **Strict 
type coercion rules:** implicit casts between incompatible types are 
disallowed. - **Standard SQL arithmetic behavior [...]
 | datafusion.execution.hash_join_buffering_capacity                       | 0  
                       | How many bytes to buffer in the probe side of hash 
joins while the build side is concurrently being built. Without this, hash 
joins will wait until the full materialization of the build side before polling 
the probe side. This is useful in scenarios where the query is not completely 
CPU bounded, allowing to do some early work concurrently and reducing the 
latency of the query. Note tha [...]


---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to