xlsongc commented on code in PR #73926:
URL: https://github.com/apache/airflow/pull/73926#discussion_r4165498006
##########
providers/common/ai/src/airflow/providers/common/ai/operators/document_loader.py:
##########
@@ -452,19 +454,56 @@ def _parse_pdf_stream(self, stream: BinaryIO) ->
list[dict[str, Any]]:
def _parse_docx_stream(self, stream: BinaryIO) -> list[dict[str, Any]]:
"""
- Parse a DOCX stream into documents.
+ Parse a DOCX stream into a single document.
- Extracts paragraph text only. Tables, headers, footers, and footnotes
- are not included. For richer DOCX parsing, plug in a dedicated
- extraction tool (``Unstructured``, ``docling``) as a custom parser
- backend.
+ Paragraphs and tables in the document body are extracted in document
+ order. Each table row becomes one "| cell | cell |" line, and a nested
+ table is flattened into its cell. Headers, footers, footnotes, and
+ content controls are not included.
"""
try:
from docx import Document
+ from docx.table import Table
except ImportError as e:
raise AirflowOptionalProviderFeatureException(e)
doc = Document(stream)
- paragraphs = [p.text for p in doc.paragraphs if p.text.strip()]
- text = "\n\n".join(paragraphs)
+ blocks = []
+ for block in doc.iter_inner_content():
+ if isinstance(block, Table):
+ text = "\n".join(f"| {' | '.join(cells)} |" for cells in
self._get_docx_table_rows(block))
Review Comment:
Agreed, done. Each top-level table is wrapped: a `ValueError` or
`RecursionError` from python-docx skips that table with a warning naming the
file, and the rest of the document still loads. I kept the catch to those two
so other errors still fail the task. One test uses a real first-row
`vMerge="continue"`; `RecursionError` is mocked, since a real one needs ~1000
rows and ~20 s. The docs mention the skip too.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]