[
https://issues.apache.org/jira/browse/FLINK-40194?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Dian Fu reassigned FLINK-40194:
-------------------------------
Assignee: Timo Theusner
> Add drop_duplicates to DataFrame API
> ------------------------------------
>
> Key: FLINK-40194
> URL: https://issues.apache.org/jira/browse/FLINK-40194
> Project: Flink
> Issue Type: Sub-task
> Components: API / Python
> Reporter: Dian Fu
> Assignee: Timo Theusner
> Priority: Major
> Fix For: 2.4.0
>
>
> def drop_duplicates(
> self,
> subset: Union[str, List[str]] = None,
> *,
> keep: str = "first",
> order_by: Union[str, Expression, List[Union[str, Expression]]] = None,
> nulls_first: Union[bool, List[bool]] = None,
> ) -> "DataFrame"
> Remove duplicate rows. When subset is omitted, fully identical rows are
> dropped; when subset is given, rows are deduplicated by those key columns,
> keeping one row per key. distinct and unique are provided as aliases of this
> method, matching the common naming convention used by other DataFrame
> libraries.
> * subset — a column name or list of column names that define a duplicate.
> When omitted (default), all columns are considered (whole-row distinct).
> * keep — "first" keeps the earliest row, "last" keeps the latest row within
> each duplicate group.
> * order_by — a column name or Expression (or a list of them) that defines the
> order in which keep selects the surviving row. When omitted, processing time
> is used, so keep="first"/"last" keep the first/last row to arrive. Ignored
> when subset is omitted.
> * nulls_first — a bool, or a list of bool aligned with order_by, controlling
> where NULLs rank when keep selects the surviving row. When omitted, Flink's
> default null ordering applies. Ignored when order_by is omitted.
> Example:
> {code:python}
> # Whole-row deduplication: drop fully identical rows
> df.drop_duplicates()
> # Keep the earliest row per device by arrival (defaults to processing time)
> df.drop_duplicates(subset="device_id")
> # Keep the earliest event per order, ordered by event time
> df.drop_duplicates(subset=["order_id"], order_by="event_time")
> # Keep the latest record per user, ordered by event time
> df.drop_duplicates(subset="user_id", order_by="event_time", keep="last")
> {code}
--
This message was sent by Atlassian Jira
(v8.20.10#820010)