ycbowal opened a new pull request, #74177: URL: https://github.com/apache/airflow/pull/74177
closes: #72592 ## Description `BedrockCreateKnowledgeBaseOperator` could only create self-managed VECTOR knowledge bases: the `knowledgeBaseConfiguration` was hardcoded, `storageConfiguration` was always sent, and overriding either through `create_knowledge_base_kwargs` raised a `TypeError`. This PR adds a `knowledge_base_config` parameter, passed as-is to the API, which supports MANAGED knowledge bases as well as any other type (KENDRA, SQL, and future ones). ## Changes - New templated `knowledge_base_config` parameter, mutually exclusive with `embedding_model_arn`. - `embedding_model_arn` and `storage_config` are now optional. Without `knowledge_base_config`, the operator still builds the VECTOR configuration from `embedding_model_arn`, so existing DAGs are unaffected. - `storageConfiguration` is only sent when `storage_config` is provided. - The vector index retry loop (`wait_for_indexing`) only applies when `storage_config` is provided. - Keys covered by a dedicated parameter (`name`, `roleArn`, `knowledgeBaseConfiguration`, `storageConfiguration`) are rejected in `create_knowledge_base_kwargs` with a clear `ValueError`. They previously raised a `TypeError`, so no existing DAG is affected. - Documentation updated with a managed knowledge base example and a link to the API structure. ## Design notes - The mutual exclusivity of `knowledge_base_config` and `embedding_model_arn` is checked in the constructor with `is None` / `is not None`, following the "Templated fields in Operator's __init__ method" section of the contributing docs, since it only checks whether arguments were passed. Value checks remain in `execute()`. - I chose to raise when both are provided rather than silently prefer one, to avoid ambiguity about which embedding model is used (e.g. a MANAGED base with a CUSTOM embedding model, where the ARN belongs inside the configuration). Happy to switch to precedence with a warning if maintainers prefer. - The knowledge base type is intentionally not validated, so new types added by AWS work without a provider release. - Note: the issue mentions `embeddingModeType`; the actual API field is `embeddingModelType`. ## Tests - Creating an operator with `knowledge_base_config` only, for every non-vector type (MANAGED with managed and custom embedding models, KENDRA, SQL), sends the configuration as-is without `storageConfiguration`. - `knowledge_base_config` is read at `execute()` time, after template rendering. - Invalid combinations (both or neither of `knowledge_base_config` and `embedding_model_arn`) are rejected at construction. - The index retry loop is skipped when no `storage_config` is provided. - Dedicated keys are rejected in `create_knowledge_base_kwargs`, and other keys are passed through. Existing tests pass unchanged, confirming backward compatibility. The existing system test (`example_bedrock_retrieve_and_generate.py`) is also unchanged and still covers the vector path. The managed knowledge base is documented with a code example; I can add a system test for it if maintainers find it useful. --- ##### Was generative AI tooling used to co-author this PR? - [x] Yes Generated-by: Claude (Anthropic), used as a step-by-step assistant for design discussion, code and test review, following the [Gen-AI guidelines](https://github.com/apache/airflow/blob/main/contributing-docs/05_pull_requests.rst#gen-ai-assisted-contributions). All changes were reviewed, understood and tested locally. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
