Messages by Thread
-
-
Re: [I] `corr(DISTINCT x, x)` fails with internal error in `SingleDistinctToGroupBy` [datafusion]
via GitHub
-
Re: [I] Substrait producer emits LIKE with three arguments and an undefined ilike function [datafusion]
via GitHub
-
Re: [PR] fix: array_position nested lists handling & check index for nulls [datafusion]
via GitHub
-
Re: [PR] Decouple the `FileSource` trait from the concrete `FileScanConfig` `DataSource` [datafusion]
via GitHub
-
Re: [PR] feat(spark): add levenshtein with optional threshold support [datafusion]
via GitHub
-
Re: [PR] benchmarks: count simulated object store GET requests per query [datafusion]
via GitHub
-
Re: [PR] fix: correlated exists/not exists subqueries hit the count bug for groupless aggregates [datafusion]
via GitHub
-
Re: [I] `RewriteSetComparison` is overly aggressive rewriting some set comparison into multiple mark joins [datafusion]
via GitHub
-
[I] percentile of -0.0 returns 0.0 [datafusion-comet]
via GitHub
-
Re: [I] Improve NDV statistics for Parquet columns (parquet RLE & Dictionary array focus) [datafusion]
via GitHub
-
Re: [I] Return `NativeType` instead of `DataType` for `get_example_types` [datafusion]
via GitHub
-
[I] Joins: one shared protocol for "the last probe partition emits the build-side rows" [datafusion]
via GitHub
-
Re: [PR] fix: evaluate sort-merge join filters that reference no columns [datafusion]
via GitHub
-
Re: [PR] feat: cost-based join order enumeration [datafusion]
via GitHub
-
Re: [PR] docs: document casts involving arrays, structs, and maps [datafusion-comet]
via GitHub
-
Re: [PR] ci: add performance regression check for PRs / merges [datafusion]
via GitHub
-
[PR] feat: support releasing batches in RecordBatchMemoryCounter [datafusion]
via GitHub
-
Re: [PR] fix: decorrelate grouping sets that leave out the correlated column when safe [datafusion]
via GitHub
-
[PR] fix: avoid sort-based joins for nested float join keys with negative zeros [datafusion]
via GitHub
-
[PR] feat: report the sort order of sorted Iceberg scans to DataFusion [datafusion-iceberg]
via GitHub
-
[PR] fix: don't let S3A committer, delete and read defaults block native Iceberg writes [datafusion-comet]
via GitHub
-
[PR] fix: preserve common field metadata through CASE expressions [datafusion]
via GitHub
-
Re: [I] `CASE` expressions drop field metadata, silently stripping Arrow extension types and breaking metadata-aware UDFs [datafusion]
via GitHub
-
[I] perf: every Comet broadcast hash join probe task re-decodes the whole broadcast in the JVM [datafusion-comet]
via GitHub
-
[I] perf: without AQE, native operators read Comet shuffle through a JVM round trip per block, making small reduce tasks up to 4x slower than Spark [datafusion-comet]
via GitHub
-
[I] Leaf expression pushdown leaves an unqualified column in an Aggregate's output name, so the plan fails a datafusion-proto round trip [datafusion]
via GitHub
-
[I] Agree on the scope and repositories for Apache Comet [datafusion-comet]
via GitHub
-
[I] [EPIC] Promote Comet to a top level Apache project [datafusion-comet]
via GitHub
-
Re: [PR] docs: add version picker and manual release snapshots [datafusion]
via GitHub
-
Re: [I] [Feature] Support Spark expression: subtract_dates [datafusion-comet]
via GitHub
-
[PR] build(deps): bump rustls from 0.23.40 to 0.23.45 [datafusion-python]
via GitHub
-
[I] Native hash joins lose the streamed side's order when it arrives sorted from outside the native plan [datafusion-comet]
via GitHub
-
Re: [PR] build(deps): bump soupsieve from 2.8.4 to 2.9 [datafusion-python]
via GitHub
-
Re: [PR] Let extension bundles declare scalar, aggregate, and window functions [datafusion-python]
via GitHub
-
[I] ShuffleScan metrics are dropped under AQE: shuffle read and decode time missing on the direct-read path [datafusion-comet]
via GitHub
-
Re: [PR] fix: Estimate string equality and LIKE filter selectivity [datafusion]
via GitHub
-
Re: [PR] POC: batch-granular Parquet scans with the push decoder and scan_plan read-ahead [datafusion]
via GitHub
-
Re: [PR] [asf-site] docs: publish 55.0.0 and 55.1.0 documentation snapshots [datafusion]
via GitHub
-
Re: [PR] docs: add datafusion-arrowmetal to the integrations list [datafusion]
via GitHub
-
[PR] perf: read shuffle blocks in 64 KiB pieces instead of through `Channels.newChannel` [datafusion-comet]
via GitHub
-
Re: [PR] docs: preserve released documentation during deployment [datafusion]
via GitHub
-
[PR] WIP speed up hashes [datafusion]
via GitHub
-
Re: [PR] chore: Hide internal public utility APIs [datafusion]
via GitHub
-
[I] Math functions differ from Spark in the last digits [datafusion-comet]
via GitHub
-
[I] Null-named attribute fails Arrow export with "field name cannot be null" [datafusion-comet]
via GitHub
-
[I] from_csv: wrong results in permissive modes and failure with a variant schema [datafusion-comet]
via GitHub
-
Re: [PR] build(deps): bump pyjwt from 2.13.0 to 2.15.0 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump github/codeql-action/init from 4.37.9 to 4.38.2 [datafusion-python]
via GitHub
-
[I] CometIcebergWriteDetectionSuite assumes a split Iceberg write plan [datafusion-comet]
via GitHub
-
[I] collect_set / collect_list return elements in a different order than Spark [datafusion-comet]
via GitHub
-
Re: [PR] build(deps): bump virtualenv from 21.5.0 to 21.7.13 [datafusion-python]
via GitHub
-
Re: [I] Enable spark.comet.exec.localTableScan.enabled when running Spark SQL tests [datafusion-comet]
via GitHub
-
[I] GROUP BY on a struct containing a map fails with CometNativeException [datafusion-comet]
via GitHub
-
[I] Native errors surface as SparkException where Spark raises a specific error class [datafusion-comet]
via GitHub
-
[I] test: Spark SQL tests that inspect the physical plan fail when LocalTableScanExec is native [datafusion-comet]
via GitHub
-
Re: [PR] build(deps): bump urllib3 from 2.7.0 to 2.8.0 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump astral-sh/setup-uv from 10.0.1 to 10.2.0 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump github/codeql-action/analyze from 4.37.9 to 4.38.2 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump tornado from 6.5.8 to 6.5.9 [datafusion-python]
via GitHub
-
Re: [I] Move CometCollationSuite into the spark-4.1+ test shim to avoid per-version duplication [datafusion-comet]
via GitHub
-
Re: [PR] chore: move CometCollationSuite to spark-4.1+ test shim [datafusion-comet]
via GitHub
-
[I] perf: divide once per row in `pmod` and `div_round_half_up` [datafusion-comet]
via GitHub
-
[PR] BigQuery: Allow tokens after a hyphenated identifier ending in a number [datafusion-sqlparser-rs]
via GitHub
-
Re: [PR] docs: add local-diff mode to review-comet-pr skill and ask about it in the PR template [datafusion-comet]
via GitHub
-
[I] [EPIC] Set up the Apache Comet project after the board adopts the resolution [datafusion-comet]
via GitHub
-
Re: [I] Support fs.s3a.auth.profile.name and fs.s3a.auth.profile.file for ProfileCredentialsProvider [datafusion-comet]
via GitHub
-
Re: [I] Native Azure store lets ambient AZURE_* environment variables override or corrupt explicit Hadoop auth config [datafusion-comet]
via GitHub
-
[I] Native Iceberg scan returns NULL for nested fields of migrated tables that use a name mapping [datafusion-comet]
via GitHub
-
Re: [I] fair_unified: a JVM consumer freeing its last bytes can fail a parked native acquire with NoSuchElementException [datafusion-comet]
via GitHub
-
[I] Native Iceberg scan returns nested values that Spark reads as NULL in files without field ids or a name mapping [datafusion-comet]
via GitHub
-
[PR] Use avro for datafiles collection from write-exec [datafusion-iceberg]
via GitHub
-
[I] Draft the Apache Comet proposal and board resolution [datafusion-comet]
via GitHub
-
[I] Agree on the initial PMC, committers, and chair for Apache Comet [datafusion-comet]
via GitHub
-
Re: [I] A JVM consumer's parked page allocation fails when another consumer of the task empties its balance [datafusion-comet]
via GitHub
-
Re: [I] virtual_hosted_style_request bad calculation [datafusion-comet]
via GitHub
-
[PR] feat: run broadcast joins natively when an unsupported scan is on the build side [datafusion-comet]
via GitHub
-
[PR] feat: txt file support [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: timestamp_add_ym_interval [datafusion-comet]
via GitHub
-
Re: [PR] fix: honor Spark’s commit protocol in Spark 3.x native writes [datafusion-comet]
via GitHub
-
Re: [PR] fix: restore localReads assertion in AdaptiveQueryExecSuite diffs [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: timestamp_add_interval [datafusion-comet]
via GitHub
-
[PR] fix: choose the forced shuffled hash join at planning time so Spark keeps its sorts [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: subtract_timestamps [datafusion-comet]
via GitHub
-
Re: [I] Enable SPARK-57298 collect_set tests when Spark 4.2 SQL diff lands [datafusion-comet]
via GitHub
-
Re: [I] Iceberg serde: delete-file fields fall back to wrong defaults on reflection failure [datafusion-comet]
via GitHub
-
Re: [I] [Feature] Support Spark expression: date_add_interval [datafusion-comet]
via GitHub
-
Re: [PR] build(deps): bump arrow-select from 59.3.0 to 60.0.0 [datafusion-python]
via GitHub
-
Re: [PR] build(deps): bump arrow from 59.3.0 to 60.0.0 [datafusion-python]
via GitHub