vbhanuchander-lang opened a new pull request, #9038: URL: https://github.com/apache/devlake/pull/9038
Closes #8840 ### Summary [GitHub issue fields](https://github.blog/changelog/2026-07-02-issue-fields-are-now-generally-available/) are organization-level structured issue metadata — typed, mutually exclusive within a field, and shared across every repository in the org. They went **generally available on 2026-07-02** (the issue was filed while they were still in preview). That is precisely the gap @oliviertassinari described in #8840: `type:` labels are not mutually exclusive (`type: bug` and `type: regression` can both be applied) and are repo-scoped, so cross-org reporting is unreliable. This PR collects issue field values and lets a scope config map a field onto an issue column, where it takes precedence over the existing label regexes. ### What's included **New table `_tool_github_issue_field_values`** — one row per issue per field. `value` holds a queryable text form regardless of the field's type, and `raw_value` keeps the original JSON so anything that doesn't fit the text form (a `multi_select` array) is still recoverable: | data_type | `value` | |---|---| | `single_select` | the selected option name | | `multi_select` | option names, comma separated, in API order | | `text` | the text | | `number` | formatted without a trailing `.0` when integral, so an effort of 5 reads as `5` | | `date` | as returned (`YYYY-MM-DD`) | **New subtasks** `Collect Issue Field Values` and `Extract Issue Field Values`, hitting `GET /repos/{owner}/{repo}/issues/{issue_number}/issue-field-values`. Both are `EnabledByDefault: false`, so existing pipelines are untouched until someone opts in. **New scope config keys**, each holding a field *name*: `issueFieldPriority`, `issueFieldSeverity`, `issueFieldComponent`, `issueFieldStoryPoint`, `issueFieldDueDate` When a mapping is set and the issue has a value for that field, it wins over the label regex. Field names are matched case-insensitively. `Effort` → `story_point` and `Target date` → `due_date` line up with the four fields GitHub preconfigures out of the box (Priority, Effort, Start date, Target date). ### Two design decisions worth your attention **1. The mapping is applied in the convertor, not written back to `_tool_github_issues`.** My first version added an `Enrich Issue Fields` subtask writing into `_tool_github_issues`, mirroring how the label regexes work today in `issue_extractor.go`. The table sorter rejected it, correctly: ``` cyclic dependency detected: map[Collect Issue Field Values:[Enrich Issue Fields] ... Enrich Issue Fields:[Extract Issue Field Values Enrich Issue Fields] ...] ``` The collector iterates `_tool_github_issues` to build its request URLs, so anything that both reads and writes that table closes a loop. Applying the mapping in `ConvertIssues` removes the cycle, keeps the tool layer as raw GitHub truth, and avoids adding columns to `_tool_github_issues` at all. `fieldMapping.asSubtaskConfig()` is passed as `SubtaskConfig` so changing or clearing a mapping forces a full re-run rather than being skipped incrementally. **2. A 404 is treated as "no field values", not an error.** An org that has never configured issue fields, or a token without visibility of them, gets 404 rather than an empty list. Failing the task there would mean one such repository breaks an otherwise fine collection — the same failure shape as #9026. ### Performance `loadIssueFieldValues` fetches the mapped fields for the connection in **one** query and indexes them by issue id, rather than one query per issue. Bounded by (issues × mapped fields), so at most 5 rows per issue. Collection itself is one request per issue, which is inherent to the endpoint — it has no repo-wide variant. If you would prefer this go through `github_graphql` instead to batch it, say so and I will move it. ### Tests 14 unit tests, all passing: - `issue_field_value_extractor_test.go` — `displayValue` across all five data types using the documented response bodies, including integral vs fractional numbers, `null`, and `multi_select` arriving without option objects; `scalarValue` on missing/unknown shapes. - `issue_field_mapping_test.go` — mapping resolution and trimming, nil scope config, name lowercasing and dedup (one field driving two columns), case-insensitive and nil-safe lookup with no cross-issue leakage, date parsing across the documented and timestamp forms. ``` ok github.com/apache/incubator-devlake/plugins/github/tasks ok github.com/apache/incubator-devlake/plugins/github/impl <- subtask graph is acyclic ok github.com/apache/incubator-devlake/plugins/github/models ``` `migration-script-lint` passes (the script uses `archived.NoPKModel`, not `core/models/common`). **I have not added an e2e dataflow test.** `NewDataFlowTester` panics without `E2E_DB_URL` and I have no DB here, so I would be shipping a test I could not run — every existing test in `plugins/github/e2e` fails the same way locally. Happy to add one with raw/snapshot CSV fixtures if you would like it; I would just be relying on CI to tell me it is right. ### Follow-ups - Docs for the new scope config keys belong in `devlake-website`. Happy to send that PR alongside. - Org-level field definitions (`GET /orgs/{org}/issue-fields`) are deliberately not collected: the value endpoint already returns `issue_field_name` and `data_type`, and skipping it avoids requiring org-level permissions. Easy to add if you want field descriptions or option colours in their own table. - config-ui has no inputs for the new keys yet; they are settable through the API. Let me know if you want the UI in this PR or a follow-up. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
