Hi, Weiqing.
Thanks for driving this FLIP forward again, +1
--
Best!
Xuyang
在 2026-07-17 14:21:21,"Weiqing Yang" <[email protected]> 写道:
>Hi all,
>
>Gentle bump on this.
>
>The proposal is unchanged at a high level. It's been re-aligned with the
>current master, and the earlier overhead concern is addressed (details in
>my previous message).
>
>Shengkai, Alan, Xuyang, Zakelly — thank you for the thorough review last
>time. Yuepeng and Shengkai, thank you for your votes on the earlier round.
>Since you already know this proposal well, I'd really appreciate a quick
>check that the points you raised still sit right in the current version.
>
>If there are no other concerns over the next few days, I'll start a fresh
>[VOTE] thread so the votes reflect the current design.
>
>FLIP doc:
>https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit
>
>Best,
>Weiqing
>
>On Thu, Jul 9, 2026 at 7:20 PM Weiqing Yang <[email protected]>
>wrote:
>
>> Hi all,
>>
>> I'd like to revive this discussion. This FLIP was originally proposed and
>> went to a vote last year (it received two +1s) [1]. Since the original
>> proposal and vote were about a year ago, I want to revisit the discussion
>> before starting a fresh vote to ensure the design remains fully aligned
>> with the current master branch.
>>
>> At a high level, the proposal is unchanged from the previous round. The
>> one substantive refinement since the earlier votes is the sampling approach
>> detailed below, which directly addresses the performance overhead question
>> Zakelly raised (the last open item on the thread).
>>
>> Overhead (Zakelly's question): Processing time is no longer measured on
>> every invocation. It now uses the same counter-based sampling as Flink's
>> state latency tracking (FLINK-21736). A new
>> `table.exec.udf-metric.sample-interval` option (default 100) means only
>> every Nth invocation is timed. The non-sampled fast path is a single
>> integer increment, and `udfProcessingTime` is now a Histogram
>> (p50/p75/p95/p99) backed by a bounded 128-entry circular buffer. Combined
>> with the existing `table.exec.udf-metric-enabled` gate (off by default,
>> meaning nothing is registered when disabled), this provides two robust
>> layers of protection against performance overhead.
>>
>> - Proposal doc (updated): [2]
>>
>> - cwiki FLIP-485: [3]
>>
>> - Draft PR: [4] (implements the proposal; kept in draft until we converge)
>>
>> I welcome your renewed feedback and any questions. If there are no further
>> concerns, I'll start a fresh [VOTE] thread on the updated proposal so the
>> final votes reflect the current design.
>>
>> Thanks,
>>
>> Weiqing
>>
>> [1] Previous vote thread:
>> https://lists.apache.org/thread/d0sv36839p5h03t3okv89pco2jy6vbg3
>>
>> [2]
>> https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit
>>
>> [3]
>> https://cwiki.apache.org/confluence/spaces/FLINK/pages/373885706/FLIP-485+Add+UDF+Metrics
>>
>> [4] https://github.com/apache/flink/pull/28692
>>
>>
>> On Wed, Jun 3, 2026 at 4:59 PM Weiqing Yang <[email protected]>
>> wrote:
>>
>>> Hi Zakelly,
>>>
>>> Kindly pinging here to see if you had any remaining concerns regarding
>>> the FLIP.
>>>
>>> If there are no further questions or concerns from anyone, I plan to
>>> close this discussion thread and proceed with the vote thread.
>>>
>>> Thanks,
>>> Weiqing
>>>
>>> On Tue, Mar 3, 2026 at 2:58 PM Weiqing Yang <[email protected]>
>>> wrote:
>>>
>>>> Hi Zakelly,
>>>>
>>>>
>>>> Thanks for the feedback and sorry for the late response - I am now
>>>> picking it back up.
>>>>
>>>> You raised a great point about the performance overhead, referencing
>>>> FLINK-16444 <https://issues.apache.org/jira/browse/FLINK-16444>. I've
>>>> updated the FLIP to adopt the same counter-based sampling approach used by
>>>> Flink's state latency tracking (FLINK-21736
>>>> <https://issues.apache.org/jira/browse/FLINK-21736>). Specifically:
>>>>
>>>> 1. New config: table.exec.udf-metric.sample-interval (default: 100
>>>> [1]) - only every Nth invocation is measured
>>>> 2. Fast path: Non-sampled invocations are a single integer increment -
>>>> negligible overhead
>>>> 3. Sampled path: System.nanoTime() around the UDF call, stored in a
>>>> DescriptiveStatisticsHistogram with a bounded 128-entry circular buffer [2]
>>>> 4. Metric type change: udfProcessingTime is now a Histogram (reports
>>>> p50/p75/p95/p99/mean/min/max) instead of the original Gauge
>>>> 5. Exception counting: Not sampled, since exceptions are rare events
>>>> and counting each one has negligible cost
>>>>
>>>> Combined with the existing feature gate (table.exec.udf-metric-enabled
>>>> defaulting to false), users have two layers of protection: the feature is
>>>> off by default, and when enabled, sampling keeps overhead minimal.
>>>> The updated FLIP is here: link
>>>> <https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit?tab=t.0#heading=h.ljww281maxj1>
>>>>
>>>> Would this address your concern? If so, it would be great to have your
>>>> vote on the vote thread [3].
>>>>
>>>> [1] 100: state.latency-track.sample-interval default value
>>>>
>>>> [2] 128: state.latency-track.history-size default value (line 55), which
>>>> is the circular buffer size for the DescriptiveStatisticsHistogram
>>>> [3] https://lists.apache.org/thread/d0sv36839p5h03t3okv89pco2jy6vbg3
>>>>
>>>> Thanks,
>>>> Weiqing
>>>>
>>>> On Thu, Aug 21, 2025 at 12:24 AM Zakelly Lan <[email protected]>
>>>> wrote:
>>>>
>>>>> Hi Weiqing,
>>>>>
>>>>> Sorry for the late reply. And I have one question:
>>>>>
>>>>> I'm wondering whether the UDF processing time is measured for every
>>>>> individual UDF invocation, with the average then reported, or if
>>>>> sampling
>>>>> is used instead? I'm concerned about the potential overhead if we
>>>>> measure
>>>>> every single invocation. We've encountered similar performance issues
>>>>> when
>>>>> implementing state latency tracking [1].
>>>>>
>>>>>
>>>>> [1] https://issues.apache.org/jira/browse/FLINK-16444
>>>>>
>>>>> Best,
>>>>> Zakelly
>>>>>
>>>>> On Fri, Aug 15, 2025 at 5:04 AM Weiqing Yang <[email protected]>
>>>>> wrote:
>>>>>
>>>>> > Cool - I’ll proceed to start the VOTE.
>>>>> > Thanks!
>>>>> >
>>>>> > Weiqing
>>>>> >
>>>>> > On Thu, Aug 14, 2025 at 12:53 AM Shengkai Fang <[email protected]>
>>>>> wrote:
>>>>> >
>>>>> > > I don't have any more comments.
>>>>> > >
>>>>> > > Best,
>>>>> > > Shengkai
>>>>> > >
>>>>> > > Weiqing Yang <[email protected]> 于2025年8月14日周四 14:47写道:
>>>>> > >
>>>>> > > > Thanks, Shengkai. I’ve updated the proposal doc with the
>>>>> recommended
>>>>> > > > configuration name. Please let me know if you have any additional
>>>>> > > feedback.
>>>>> > > >
>>>>> > > > Best,
>>>>> > > > Weiqing
>>>>> > > >
>>>>> > > > On Wed, Aug 13, 2025 at 6:58 PM Shengkai Fang <[email protected]>
>>>>> > wrote:
>>>>> > > >
>>>>> > > > > Sorry for the late response. I prefer to use
>>>>> > > > > `table.exec.udf-metric-enabled` as the option name.
>>>>> > > > >
>>>>> > > > > Best,
>>>>> > > > > Shengkai
>>>>> > > > >
>>>>> > > > > Weiqing Yang <[email protected]> 于2025年8月13日周三 23:54写道:
>>>>> > > > >
>>>>> > > > > > Hi Shengkai, Alan, Xuyang, and all,
>>>>> > > > > >
>>>>> > > > > > Since there have been no further objections, I’ll proceed to
>>>>> start
>>>>> > > the
>>>>> > > > > VOTE
>>>>> > > > > > on this proposal shortly.
>>>>> > > > > >
>>>>> > > > > > Thanks,
>>>>> > > > > > Weiqing
>>>>> > > > > >
>>>>> > > > > > On Thu, Jul 31, 2025 at 10:26 PM Weiqing Yang <
>>>>> > > > [email protected]>
>>>>> > > > > > wrote:
>>>>> > > > > >
>>>>> > > > > > > Hi Shengkai, Alan and Xuyang,
>>>>> > > > > > >
>>>>> > > > > > > Just checking in - do you have any concerns or feedback?
>>>>> > > > > > >
>>>>> > > > > > > If there are no further objections from anyone, I’ll mark
>>>>> the
>>>>> > FLIP
>>>>> > > as
>>>>> > > > > > > ready for voting.
>>>>> > > > > > >
>>>>> > > > > > >
>>>>> > > > > > > Best,
>>>>> > > > > > > Weiqing
>>>>> > > > > > >
>>>>> > > > > > >
>>>>> > > > > > > On Mon, Jul 14, 2025 at 9:10 PM Weiqing Yang <
>>>>> > > > [email protected]
>>>>> > > > > >
>>>>> > > > > > > wrote:
>>>>> > > > > > >
>>>>> > > > > > >> Hi Xuyang,
>>>>> > > > > > >>
>>>>> > > > > > >> Thank you for reviewing the proposal!
>>>>> > > > > > >>
>>>>> > > > > > >> I’m planning to use: *udf.metrics.process-time* and
>>>>> > > > > > >> *udf.metrics.exception-count*. These follow the naming
>>>>> > convention
>>>>> > > > used
>>>>> > > > > > >> in Flink (e.g., RocksDB native metrics
>>>>> > > > > > >> <
>>>>> > > > > >
>>>>> > > > >
>>>>> > > >
>>>>> > >
>>>>> >
>>>>> https://nightlies.apache.org/flink/flink-docs-master/docs/deployment/config/#rocksdb-native-metrics
>>>>> > > > > > >).
>>>>> > > > > > >> I’ve added these names to the proposal doc.
>>>>> > > > > > >>
>>>>> > > > > > >> Alternatively, I also considered:
>>>>> > > *metrics.udf.process-time.enabled*
>>>>> > > > > and
>>>>> > > > > > >> *metrics.udf.exception-count.enabled. *
>>>>> > > > > > >>
>>>>> > > > > > >> Happy to hear any feedback on which style might be more
>>>>> > > appropriate.
>>>>> > > > > > >>
>>>>> > > > > > >>
>>>>> > > > > > >> Best,
>>>>> > > > > > >> Weiqing
>>>>> > > > > > >>
>>>>> > > > > > >> On Mon, Jul 14, 2025 at 2:55 AM Xuyang <[email protected]
>>>>> >
>>>>> > > wrote:
>>>>> > > > > > >>
>>>>> > > > > > >>> Hi, Weiqing.
>>>>> > > > > > >>>
>>>>> > > > > > >>> Thanks for driving to improve this. I just have one
>>>>> question. I
>>>>> > > > > notice
>>>>> > > > > > a
>>>>> > > > > > >>> new configuration is introduced in this flip. I just
>>>>> wonder
>>>>> > what
>>>>> > > > the
>>>>> > > > > > >>> configuration name is. Could you please include the full
>>>>> name
>>>>> > of
>>>>> > > > this
>>>>> > > > > > >>> configuration? (just similar to the other names in
>>>>> > > MetricOptions?)
>>>>> > > > > > >>>
>>>>> > > > > > >>>
>>>>> > > > > > >>>
>>>>> > > > > > >>>
>>>>> > > > > > >>> --
>>>>> > > > > > >>>
>>>>> > > > > > >>> Best!
>>>>> > > > > > >>> Xuyang
>>>>> > > > > > >>>
>>>>> > > > > > >>>
>>>>> > > > > > >>>
>>>>> > > > > > >>>
>>>>> > > > > > >>>
>>>>> > > > > > >>> 在 2025-07-13 12:03:59,"Weiqing Yang" <
>>>>> [email protected]
>>>>> > >
>>>>> > > > 写道:
>>>>> > > > > > >>> >Hi Alan,
>>>>> > > > > > >>> >
>>>>> > > > > > >>> >Thanks for reviewing the proposal and for highlighting
>>>>> the
>>>>> > > > > ASYNC_TABLE
>>>>> > > > > > >>> work.
>>>>> > > > > > >>> >
>>>>> > > > > > >>> >Yes, I’ve updated the proposal to cover both
>>>>> ASYNC_SCALAR and
>>>>> > > > > > >>> ASYNC_TABLE.
>>>>> > > > > > >>> >For async UDFs, the plan is to instrument both the
>>>>> > invokeAsync()
>>>>> > > > > call
>>>>> > > > > > >>> and
>>>>> > > > > > >>> >the async callback handler to measure the full end-to-end
>>>>> > > latency
>>>>> > > > > > until
>>>>> > > > > > >>> the
>>>>> > > > > > >>> >result or error is returned from the future.
>>>>> > > > > > >>> >
>>>>> > > > > > >>> >Let me know if you have any further questions or
>>>>> suggestions.
>>>>> > > > > > >>> >
>>>>> > > > > > >>> >Best,
>>>>> > > > > > >>> >Weiqing
>>>>> > > > > > >>> >
>>>>> > > > > > >>> >On Thu, Jul 10, 2025 at 4:15 PM Alan Sheinberg
>>>>> > > > > > >>> ><[email protected]> wrote:
>>>>> > > > > > >>> >
>>>>> > > > > > >>> >> Hi Weiqing,
>>>>> > > > > > >>> >>
>>>>> > > > > > >>> >> From your doc, the entrypoint for UDF calls in the
>>>>> codegen
>>>>> > is
>>>>> > > > > > >>> >> ExprCodeGenerator which should invoke
>>>>> > > > BridgingSqlFunctionCallGen,
>>>>> > > > > > >>> which
>>>>> > > > > > >>> >> could be instrumented with metrics. This works well
>>>>> for
>>>>> > > > > synchronous
>>>>> > > > > > >>> calls,
>>>>> > > > > > >>> >> but what about ASYNC_SCALAR and the soon to be merged
>>>>> > > > ASYNC_TABLE
>>>>> > > > > (
>>>>> > > > > > >>> >> https://github.com/apache/flink/pull/26567)? Timing
>>>>> > metrics
>>>>> > > > > would
>>>>> > > > > > >>> only
>>>>> > > > > > >>> >> account for what it takes to call invokeAsync, not for
>>>>> the
>>>>> > > > result
>>>>> > > > > to
>>>>> > > > > > >>> >> complete (with a result or error from the future
>>>>> object).
>>>>> > > > > > >>> >>
>>>>> > > > > > >>> >> There are appropriate places which can handle the async
>>>>> > > > callbacks,
>>>>> > > > > > >>> but they
>>>>> > > > > > >>> >> are in other locations. Will you be able to support
>>>>> those
>>>>> > as
>>>>> > > > > well?
>>>>> > > > > > >>> >>
>>>>> > > > > > >>> >> Thanks,
>>>>> > > > > > >>> >> Alan
>>>>> > > > > > >>> >>
>>>>> > > > > > >>> >> On Wed, Jul 9, 2025 at 7:52 PM Shengkai Fang <
>>>>> > > [email protected]
>>>>> > > > >
>>>>> > > > > > >>> wrote:
>>>>> > > > > > >>> >>
>>>>> > > > > > >>> >> > I just have some questions:
>>>>> > > > > > >>> >> >
>>>>> > > > > > >>> >> > 1. The current metrics hierarchy shows that the UDF
>>>>> metric
>>>>> > > > group
>>>>> > > > > > >>> belongs
>>>>> > > > > > >>> >> to
>>>>> > > > > > >>> >> > the TaskMetricGroup. I think it would be better for
>>>>> the
>>>>> > UDF
>>>>> > > > > metric
>>>>> > > > > > >>> group
>>>>> > > > > > >>> >> to
>>>>> > > > > > >>> >> > belong to the OperatorMetricGroup instead, because a
>>>>> UDF
>>>>> > > might
>>>>> > > > > be
>>>>> > > > > > >>> used by
>>>>> > > > > > >>> >> > multiple operators.
>>>>> > > > > > >>> >> > 2. What are the naming conventions for UDF metrics?
>>>>> Could
>>>>> > > you
>>>>> > > > > > >>> provide an
>>>>> > > > > > >>> >> > example? Do the metric name contains the UDF name?
>>>>> > > > > > >>> >> > 3. Why is the UDFExceptionCount metric introduced?
>>>>> If a
>>>>> > UDF
>>>>> > > > > throws
>>>>> > > > > > >>> an
>>>>> > > > > > >>> >> > exception, the job fails immediately. Why do we need
>>>>> to
>>>>> > > track
>>>>> > > > > this
>>>>> > > > > > >>> value?
>>>>> > > > > > >>> >> >
>>>>> > > > > > >>> >> > Best
>>>>> > > > > > >>> >> > Shengkai
>>>>> > > > > > >>> >> >
>>>>> > > > > > >>> >> >
>>>>> > > > > > >>> >> > Weiqing Yang <[email protected]> 于2025年7月9日周三
>>>>> > > 12:59写道:
>>>>> > > > > > >>> >> >
>>>>> > > > > > >>> >> > > Hi all,
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > I’d like to initiate a discussion about adding UDF
>>>>> > > metrics.
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > *Motivation*
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > User-defined functions (UDFs) are essential for
>>>>> custom
>>>>> > > logic
>>>>> > > > > in
>>>>> > > > > > >>> Flink
>>>>> > > > > > >>> >> > jobs
>>>>> > > > > > >>> >> > > but often act as black boxes, making debugging and
>>>>> > > > performance
>>>>> > > > > > >>> tuning
>>>>> > > > > > >>> >> > > difficult. When issues like high latency or
>>>>> frequent
>>>>> > > > > exceptions
>>>>> > > > > > >>> occur,
>>>>> > > > > > >>> >> > it's
>>>>> > > > > > >>> >> > > hard to pinpoint the root cause inside UDFs.
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > Flink currently lacks built-in metrics for key UDF
>>>>> > aspects
>>>>> > > > > such
>>>>> > > > > > as
>>>>> > > > > > >>> >> > > per-record processing time or exception count. This
>>>>> > limits
>>>>> > > > > > >>> >> observability
>>>>> > > > > > >>> >> > > and complicates:
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > - Debugging production issues
>>>>> > > > > > >>> >> > > - Performance tuning and resource allocation
>>>>> > > > > > >>> >> > > - Supplying reliable signals to autoscaling
>>>>> systems
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > Introducing standard, opt-in UDF metrics will
>>>>> improve
>>>>> > > > platform
>>>>> > > > > > >>> >> > > observability and overall health.
>>>>> > > > > > >>> >> > > Here’s the proposal document: Link
>>>>> > > > > > >>> >> > > <
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> >
>>>>> > > > > > >>> >>
>>>>> > > > > > >>>
>>>>> > > > > >
>>>>> > > > >
>>>>> > > >
>>>>> > >
>>>>> >
>>>>> https://docs.google.com/document/d/1ZTN_kSxTMXKyJcrtmP6I9wlZmfPkK8748_nA6EVuVA0/edit?tab=t.0#heading=h.ljww281maxj1
>>>>> > > > > > >>> >> > > >
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > Your feedback and ideas are welcome to refine this
>>>>> > > feature.
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> > > Thanks,
>>>>> > > > > > >>> >> > > Weiqing
>>>>> > > > > > >>> >> > >
>>>>> > > > > > >>> >> >
>>>>> > > > > > >>> >>
>>>>> > > > > > >>>
>>>>> > > > > > >>
>>>>> > > > > >
>>>>> > > > >
>>>>> > > >
>>>>> > >
>>>>> >
>>>>>
>>>>