GitHub user joeyutong edited a discussion: [Discussion][Observability] Replace 
Event Log with a Unified Trace Log

This proposal builds on [Recording Agent Traces in the Event Log 
(#900)](https://github.com/apache/flink-agents/discussions/900). It brings 
Event and execution logging into a unified Trace Log model while preserving the 
observation coverage and execution semantics of the existing design.

## 1. Background

Event Log currently serves two purposes: recording Events that flow between 
Actions and recording the execution lifecycle of Actions and their LLM, Parser, 
and Tool calls. Lifecycle reports are represented as synthetic Events, such as 
`_execution_finished_event`, even though they are used only for observability 
and are never routed to Actions.

Using Event for both purposes introduces ambiguity into the model and its 
configuration:

- **Event has two meanings.** The same term and data structure describe both 
objects in the programming model and records used to observe execution.
- **Logging configuration spans both models.** Event Log levels and per-type 
settings control ordinary Event logging, while execution logging requires an 
additional Trace switch. Selecting execution reports can also require users to 
know synthetic Event types such as `_execution_finished_event`.
- **Execution reports carry redundant metadata.** An execution already has an 
`executionId` and a `status`, but each lifecycle report also receives an Event 
UUID and type because it is represented as an Event.

We propose **replacing Event Log with a unified Trace Log**. Event will retain 
its meaning in the programming model: an object that Actions consume or emit. 
Trace Log will describe Event flow and execution through a common record format 
and a single set of logging controls, preserving the information currently 
available in Event Log.

## 2. Goals and Scope

### Goals

1. **Give Event and Trace distinct responsibilities.** Event belongs to the 
programming model; Trace provides a common representation for observations of 
Events and execution.
2. **Preserve existing observation coverage.** This includes Event content, 
execution status, Memory reads and writes, initial Memory snapshots, and the 
relationships between Events and the executions that produce or consume them.
3. **Unify logging configuration.** Users should be able to choose which 
records to write, how much content to retain, and where to send the output, 
with local overrides for specific Event types, Actions, or component calls.

### Scope

- The design covers four areas: the data model, runtime collection, 
configuration, and log consumption.
- Event routing, Action execution, and EventListener callbacks retain their 
existing behavior, as do recovery and result reuse within the same version.
- Trace remains best effort. Records may be missing or duplicated; complete 
execution histories, exactly-once logging, and deduplication after recovery are 
outside the scope of this proposal.
- Compatibility with previous Event JSON and historical Event Log formats is 
outside scope. Applications, log consumers, and external data must migrate to 
the new contracts.
- Migration of persisted runtime state across versions is also outside scope. 
Upgrade guidance must identify affected checkpoint and savepoint restore paths, 
as well as reuse of results persisted in ActionStateStore.

## 3. Proposed Design

### 3.1 Data Model and Contract

#### Record composition

TraceRecord replaces EventLogRecord as the unit written to the log. It contains 
a TraceContext, a timestamp, and attributes, with execution status and failure 
category where applicable. The timestamp records observation time by default; 
component reports may supply the occurrence time. Event observations and 
execution reports use this same record type.

The following comparison shows the object structures for the same failed 
execution. Nesting indicates field ownership; the next section illustrates the 
serialized JSON format.

```text
Current: EventLogRecord                         Proposed: TraceRecord
├── eventContext: EventContext                  ├── context: TraceContext
│   ├── timestamp                               │   ├── inputRunId, 
businessKey, agentName
│   └── eventType = "_execution_failed_event"   │   ├── entityType, entityName
├── traceContext: ExecutionTraceContext         │   ├── executionId, 
parentExecutionId
│   ├── inputRunId, businessKey, agentName      │   └── entityMetadata
│   ├── entityType, entityName                  ├── timestamp
│   ├── executionId, parentExecutionId          ├── status = "failed"
│   └── entityMetadata                          ├── problemCategory (when 
supplied)
└── event: Event                                └── attributes
    ├── id (generated UUID)                         ├── errorType
    ├── type = "_execution_failed_event"            └── errorMessage
    ├── upstreamEventId = null
    ├── upstreamActionName = null
    └── attributes
        ├── status = "failed"
        ├── problemCategory (when supplied)
        ├── errorType
        └── errorMessage
```

Execution identity remains in the context. The synthetic Event, its generated 
UUID, and its lifecycle type are removed. Status and failure category become 
fields on TraceRecord, while error details remain in `attributes`.

- **TraceContext** retains the field structure of `ExecutionTraceContext`. 
`inputRunId`, `businessKey`, and `agentName` retain their existing meanings. 
The changes extend the context to describe Events:
  - For Event records, `entityType` is `event` and `entityName` is the Event's 
type.
  - `executionId` and `parentExecutionId` retain their execution semantics and 
are absent from Event records. An Event record refers to its producer through 
`entityMetadata.producerExecutionId`.
  - `entityMetadata` holds Event identity and source information (`eventId`, 
`producerExecutionId`, `upstreamEventId`, and `upstreamActionName`) on Event 
records. Action records gain `triggerEventId` to identify the Event that 
triggered the execution.
- **TraceRecord** contains its TraceContext. Serialization places context 
fields such as `entityType` and `inputRunId` at the top level of the JSON 
record.
- **EventContext** remains part of the EventListener callback contract and is 
fully independent of TraceContext. There is no containment, inheritance, or 
conversion relationship between the two, and TraceRecord construction does not 
depend on EventContext.

Payload truncation applies only to `attributes`. IDs, relationship fields, and 
the top-level `problemCategory` remain intact, preserving the information 
needed to connect records and classify failures. Truncation affects only the 
serialized content; the Event delivered to user code is unchanged.

#### Representing an Event

An Event observation is a standalone TraceRecord. Its context identifies the 
Event, its attributes contain the Event payload, and its metadata links it to 
the producing Action execution when one exists.

Consider a `create_order` Action execution (`action-1`) that consumes Event 
`event-1` and emits an `OrderCreated` Event (`event-2`). The examples below use 
abbreviated IDs and omit timestamps and common metadata for brevity.

**Previous Event Log record, with tracing enabled.** The record combines the 
output Event with its producer's execution context: `entityType`, `entityName`, 
and `executionId` describe `create_order`, while the `event*` fields describe 
`OrderCreated`.

```json
{
  "inputRunId": "run-1",
  "entityType": "action",
  "entityName": "create_order",
  "executionId": "action-1",
  "eventId": "event-2",
  "eventType": "OrderCreated",
  "upstreamEventId": "event-1",
  "upstreamActionName": "create_order",
  "eventAttributes": {
    "orderId": "order-1"
  }
}
```

**Proposed TraceRecord.** The context now describes `OrderCreated` itself. Its 
relationship to the `create_order` execution is explicit in 
`producerExecutionId`.

```json
{
  "inputRunId": "run-1",
  "entityType": "event",
  "entityName": "OrderCreated",
  "entityMetadata": {
    "eventId": "event-2",
    "producerExecutionId": "action-1",
    "upstreamEventId": "event-1",
    "upstreamActionName": "create_order"
  },
  "attributes": {
    "orderId": "order-1"
  }
}
```

The Event's `type` maps to `entityName`, its `id` to `entityMetadata.eventId`, 
and its payload directly to `attributes`. This preserves the Event's identity 
and content without embedding the Event object in the record.

Event flow and execution relationships use the following fields:

| Relationship | Representation |
|---|---|
| An Action invokes an LLM, Parser, or Tool | The child execution's 
`parentExecutionId` points to the Action execution |
| An execution emits or replays an Event | The Event record's 
`entityMetadata.producerExecutionId` identifies that execution |
| An Event triggers an Action execution | The Action record's 
`entityMetadata.triggerEventId` identifies that Event |
| `create_order` consumes `event-1` and emits `event-2` | The `event-2` 
record's `entityMetadata.upstreamEventId` is `event-1`, and its 
`upstreamActionName` is `create_order` |

These relationships preserve a distinction between an Event's identity and the 
execution that produces or replays it:

- An Event has no execution lifecycle. Its record therefore omits 
`executionId`, `parentExecutionId`, and top-level `status` and 
`problemCategory`. User attributes with these names remain in `attributes` and 
are subject to payload truncation.
- `producerExecutionId` is absent when there is no producing Action execution, 
as with a root InputEvent. Framework-generated Events retain their existing 
source information even when no Action execution can be referenced.
- `eventId` identifies the Event, not a unique observation. The same Event may 
appear in multiple records, including when a saved output is replayed during 
recovery.
- On replay, `producerExecutionId` identifies the execution emitting the saved 
output at that point. This may differ from the execution that originally 
produced it.
- `executionId` retains its existing task creation and restoration semantics. A 
restart does not necessarily assign a new execution ID.

Source fields move from the programming-model Event to Trace metadata. Custom 
Event construction and reconstruction continue to use the Event's ID, type, and 
attributes.

### 3.2 Collection and Runtime Flow

Integrating TraceRecord into the runtime requires three changes: constructing 
records directly at the existing collection points, moving Event source 
information into those records, and retaining the context needed to connect 
records independently of logging configuration.

#### Previous Event Log flow

In the previous Event Log flow, records reached EventLogWriter through three 
paths:

- **Events:** EventRouter supplies the Event, EventContext, and optional 
ExecutionTraceContext before Action matching or downstream delivery. This path 
covers input, output, custom, and framework-generated Events, including Events 
with no consumers.
- **Action lifecycle:** ActionExecutionOperator reports start, completion, 
failure, and result reuse by creating lifecycle Events and passing them through 
ExecutionEventLogger.
- **Component calls:** Existing LLM, Parser, and Tool call sites report through 
ExecutionReporter and RunnerContext, which supply execution context and create 
lifecycle Events. Python reports use the existing Python-to-Java bridge.

#### Runtime changes

##### 1. Construct TraceRecords at the collection points

Each collection point will produce a TraceRecord describing the Event or 
execution it observes. Execution reports no longer need a synthetic Event to 
carry their status and content.

| Observation | Current construction | Proposed construction |
|---|---|---|
| Event | EventLogRecord combines the Event, EventContext, and optional 
execution context. | EventRouter constructs a TraceRecord with the Event's 
identity, content, and available run and source references. |
| Action or component execution | Reporting code creates a lifecycle Event and 
combines it with execution context. | Reporting code constructs a TraceRecord 
with execution context, status, and any failure details. |

All records then enter a common filtering, serialization, and output path 
governed by Trace Log configuration.

##### 2. Populate Event relationships from runtime context

In the previous Event Log model, the runtime wrote `upstreamEventId` and 
`upstreamActionName` onto an emitted Event and supplied its producer's 
execution context to the logger. Under the new model, the runtime places these 
relationships in Trace metadata:

- When creating an Action task, it records the triggering Event's ID in the 
Action's TraceContext as `entityMetadata.triggerEventId`.
- When constructing a TraceRecord for an Event emitted by that Action, it reads 
the triggering Event ID, Action name, and execution ID from the current task. 
These become `upstreamEventId`, `upstreamActionName`, and `producerExecutionId` 
in the Event record's `entityMetadata`.
- Component calls continue to use child execution contexts, retaining their 
execution IDs and parent execution references.

##### 3. Preserve context independently of record filtering

Omitting an execution record must not remove the context needed to describe its 
output Events. For example, an `OrderCreated` record still references the 
`create_order` execution through `producerExecutionId` even when that Action's 
execution records are not written.

The runtime therefore retains the required TraceContext through asynchronous 
resumption and result replay, regardless of which records the logger selects. A 
replayed Event record references the execution replaying it, following the 
identity semantics described in Section 3.1.

#### Preserved behavior

- **Collection timing and coverage:** Events are observed before Action 
matching or downstream delivery. Memory observations are collected during an 
Action and emitted as Events when it completes; the optional run-begin Event 
captures initial short-term Memory before the input's Actions execute. 
Framework Event generation conditions and payloads are unchanged.
- **Routing and callbacks:** Event routing and EventListener delivery retain 
their existing behavior. Listeners continue to receive EventContext and Event 
through a separate callback path. Execution reports do not enter Action routing 
or trigger EventListener callbacks.
- **Execution and recovery:** Action execution and component invocation retain 
their existing behavior, as do recovery and result reuse within the same 
version. Saved output Events continue to re-enter EventRouter after the Action 
reuse report.

Trace reporting and output remain best effort and do not change Action 
execution or recovery guarantees.

### 3.3 Configuration

All logging options live under `trace-log.*`. `trace-log.targets` combines 
record selection and recording detail in a list of targets. Each target pairs a 
required `scope` with an optional `detail`; preset and entity targets use the 
same structure.

#### Target configuration

`scope` accepts either a preset or an entity selector:

- `EVENT_ONLY` selects Events.
- `ALL` selects all observed entities.
- An object with required `entityType` and optional `entityName` selects an 
entity type or named entities within that type. Built-in types include `event`, 
`action`, `llm`, `parser`, and `tool`.

`detail` accepts `OFF`, `STANDARD`, or `VERBOSE`. `OFF` suppresses the entire 
matching record; it is not a scope and does not produce a record with empty 
attributes. `STANDARD` applies the configured attribute limits. `VERBOSE` skips 
STANDARD truncation. Both recording settings retain identities, relationship 
fields, entity metadata, timestamps, statuses, and problem categories in full. 
Existing typed `ChatMessage` media sanitization applies at both settings.

When `targets` is omitted, the default is `[{scope: EVENT_ONLY, detail: 
STANDARD}]`. An explicit list replaces this default: it does not implicitly 
retain Event logging. An empty list records nothing. A list containing only a 
Tool scope therefore records only matching Tools.

The following configuration retains Event logging, suppresses `DebugEvent`, and 
records the `create_order` Action at VERBOSE detail:

```yaml
agent:
  trace-log:
    targets:
      - scope: EVENT_ONLY
        detail: STANDARD
      - scope:
          entityType: event
          entityName: DebugEvent
        detail: "OFF"
      - scope:
          entityType: action
          entityName: create_order
        detail: VERBOSE
    standard:
      max-string-length: 2000
      max-array-elements: 20
      max-depth: 5
    output-type: slf4j
    pretty-print: false
```

The `agent` section is the YAML configuration entry point; Java and Python 
configuration APIs use the corresponding `trace-log.*` keys. In YAML, write 
`detail: "OFF"` with quotes to preserve its string value.

| Record | Effective behavior |
|---|---|
| `DebugEvent` | Omitted by its local OFF detail. |
| Other Events | Written at STANDARD, following the EVENT_ONLY preset. |
| The `create_order` Action | Its lifecycle observations are written at 
VERBOSE. |
| Other Actions and component calls, including calls inside `create_order` | 
Omitted unless another target selects them. Execution itself is unaffected. |

Each target applies to the record's own entity type and name. Selecting an 
Action does not automatically select its output Events or child calls. 
Event-only logs show flows through Actions that emit Events; an Action that 
emits no Event requires its own target to record its lifecycle.

#### Name matching and detail inheritance

Entity types and names are case-sensitive strings. Names match exactly unless 
they end in `.*`: `com.foo.*` matches `com.foo` and names beginning with 
`com.foo.`, but not `com.foobar.OrderCreated`. Omitting `entityName` selects 
the whole entity type. Other wildcard forms are unsupported. Matching by Agent 
name, business key, status, or problem category remains outside scope.

An omitted detail inherits a preset's effective detail, including OFF; it never 
inherits from another entity target:

| Target scope | Detail when omitted |
|---|---|
| `ALL` | STANDARD |
| `EVENT_ONLY` | The ALL detail, or STANDARD when ALL is absent |
| An Event entity scope | EVENT_ONLY, otherwise ALL, otherwise STANDARD |
| Any other entity scope | ALL, otherwise EVENT_ONLY, otherwise STANDARD |

When both presets exist, EVENT_ONLY supplies the Event default and ALL supplies 
the default for other entities. Inheritance determines detail only: an 
EVENT_ONLY preset does not by itself select Action or component records. An 
explicit Tool target can inherit its detail when EVENT_ONLY is the only preset.

With only an ALL preset at VERBOSE, an entity target without detail inherits 
VERBOSE. With only an ALL preset at OFF, an omitted local detail also resolves 
to OFF; that target must explicitly specify STANDARD or VERBOSE to enable 
recording.

#### Selection and precedence

For each record, the most specific matching scope determines its effective 
detail:

1. Exact entity type and name.
2. Namespace prefix within the entity type, with the longest prefix winning.
3. The whole entity type.
4. EVENT_ONLY, for Event records.
5. ALL.

A more specific target can disable recording under an enabled preset or enable 
recording under a preset with OFF detail. When no target matches, the record is 
omitted. Target order does not affect precedence. Duplicate scopes with 
different effective details after inheritance are rejected at startup; 
duplicates with the same effective detail are accepted.

Recording targets are parsed and validated at startup, and the effective 
targets are reported. These semantics are consistent across Java, Python, and 
YAML.

#### Output settings and attribute limits

Output settings sit directly under `trace-log`: `output-type`, `base-dir`, and 
`pretty-print`. `output-type` defaults to SLF4J. A non-empty `base-dir` selects 
FILE and takes precedence over `output-type`. When FILE is selected and 
`base-dir` is omitted, the directory is `java.io.tmpdir/flink-agents`.

The `standard` group contains `max-string-length`, `max-array-elements`, and 
`max-depth`, defaulting to 2000, 20, and 5. These limits apply only to 
attributes at STANDARD detail. A limit of 0 removes that particular limit 
without disabling logging.

Output remains JSONL by default. `pretty-print: true` writes multiline JSON 
objects. Memory Event and run-begin Event options continue to control Event 
generation independently of record selection; `event-listeners` is unchanged.

#### Configuration migration

Legacy logging keys are rejected at startup, including when mixed with new 
keys. They require explicit migration; no compatibility mapping is provided.

| Existing option | New setting |
|---|---|
| `event-log.trace.enabled` | An EVENT_ONLY or ALL preset in 
`trace-log.targets` |
| `event-log.level` | The preset target's `detail` |
| `event-log.type.<EVENT_TYPE>.level` | An entity target with `scope: 
{entityType: event, entityName: ...}` and `detail` |
| `event-log.standard.max-string-length` | 
`trace-log.standard.max-string-length` |
| `event-log.standard.max-array-elements` | 
`trace-log.standard.max-array-elements` |
| `event-log.standard.max-depth` | `trace-log.standard.max-depth` |
| `eventLoggerType` | `trace-log.output-type` |
| `baseLogDir` | `trace-log.base-dir` |
| `prettyPrint` | `trace-log.pretty-print` |

- Migrate `event-log.trace.enabled: false` to an EVENT_ONLY preset, or `true` 
to an ALL preset, and put the previous global level in that preset's detail. 
Migrate limits and output settings at the same time.
- Migrate ordinary Event-type settings to entity targets. Use an explicit `.*` 
namespace selector where the old configuration relied on dot-separated 
hierarchy fallback.
- Synthetic lifecycle filters have no general equivalent. A setting that 
suppressed only `_execution_finished_event` cannot be reproduced by an Action 
or Tool target, which applies to all lifecycle statuses for that entity. Status 
matching remains deferred.

Rejecting legacy keys prevents a configuration that previously disabled logging 
from silently falling back to new defaults and emitting records.

### 3.4 Output and Consumption

Both built-in output destinations serialize the same TraceRecord fields and add 
the resolved `detail` (STANDARD or VERBOSE) to the JSON. `detail` is output 
metadata, not a field on TraceRecord; OFF records are omitted. SLF4J also 
includes `jobId`, `taskName`, and `subtaskId` in each record, while file output 
identifies them in the file path. Trace records are emitted through SLF4J at 
INFO, independently of recording detail.

#### Field mappings for readers and queries

Queries and parsers must adopt the new field mappings:

| Information | Previous representation | New representation |
|---|---|---|
| Event identity and content | `eventType`, `eventId`, and `eventAttributes` | 
For `entityType = "event"`, the type is `entityName`, the ID is 
`entityMetadata.eventId`, and the content is `attributes` |
| An output Event's source | Top-level `upstreamEventId` and 
`upstreamActionName` | The same fields in `entityMetadata` |
| Execution progress | Execution fields, synthetic lifecycle Event types, and 
status | `entityType`, `entityName`, and `executionId` describe the execution; 
`status` describes its lifecycle state |
| Recording detail | Output-layer `logLevel` | Output-layer `detail` |

#### Built-in Trace Tree support

The Trace Tree tool reads the new TraceRecord format and builds an Event–Action 
graph. Execution records remain outside graph construction; this proposal does 
not add execution-state or component-call visualization.

- The reader supports only TraceRecord. Historical flat Event Log records and 
records with a nested `event` object are unsupported; no historical-format 
conversion is included.
- Event records are selected by `entityType = "event"`. An Event's name does 
not make it an execution observation.
- The graph's output structure remains unchanged.
- Observations with the same Event ID and type represent the same Event. 
Attribute differences, including STANDARD truncation, do not change that 
identity.
- Reconstruction follows Event identity and source relationships. Missing or 
conflicting relationships are reported; the reader does not invent missing 
execution IDs or relationships. Record selection may make the reconstructed 
flow incomplete.

## 4. Compatibility and Migration Boundaries

The migration affects log producers, readers, configuration, and some runtime 
structures. The following boundaries distinguish changes to observability from 
the contracts retained by the programming model.

| Interface or stored data | Compatibility boundary |
|---|---|
| EventListener | Callback signatures, timing, and EventContext behavior are 
unchanged. Trace settings do not affect callback delivery or the Event payload 
received by listeners. |
| Custom Events | Construction and reconstruction retain the Event's ID, type, 
and attributes. Code that directly accesses the removed upstream fields must be 
updated, including EventListener implementations that read those fields. |
| Event JSON | The Event model no longer includes `upstreamEventId` or 
`upstreamActionName`; source relationships belong to Trace metadata. 
Compatibility with Event JSON from previous versions is not guaranteed. 
Producers, consumers, and external data must adopt the current Event schema. |
| Internal reporting helpers | ExecutionTraceContext, lifecycle Event 
factories, and reporting internals can be refactored without a separate 
deprecation period for each internal helper. |
| Custom log output | Custom output integrations must adapt to TraceRecord and 
the new recording-detail semantics. Compatibility with the previous logging 
extension interfaces is outside scope. |
| Log format | Producers and the built-in reader use TraceRecord. Historical 
Event Log formats are unsupported. External queries and parsers must migrate 
field mappings and output metadata. |
| Operational naming | Loggers, log files, and related metrics adopt Trace 
naming. Existing log collection and monitoring configurations must be updated 
accordingly. |
| Configuration | Legacy logging keys are rejected at startup with migration 
guidance, as described in Section 3.3. |
| Persisted runtime state | Checkpoints and savepoints may contain the previous 
ActionTask and Event structures; ActionStateStore also persists triggering and 
output Events for result reuse. This proposal does not introduce cross-version 
migration for those structures. Upgrade guidance must identify affected restore 
and result-reuse paths; recovery and result reuse within the same version 
retain their existing behavior. |

Trace remains a best-effort account of runtime activity. A missing record does 
not establish that an Event or execution never occurred, and repeated records 
with the same Event ID may describe repeated observations of that Event. These 
limits apply to the logs; business execution and recovery retain their existing 
guarantees.


GitHub link: https://github.com/apache/flink-agents/discussions/1146

----
This is an automatically sent email for [email protected].
To unsubscribe, please send an email to: [email protected]

Reply via email to