[ 
https://issues.apache.org/jira/browse/BEAM-14383?focusedWorklogId=768586&page=com.atlassian.jira.plugin.system.issuetabpanels:worklog-tabpanel#worklog-768586
 ]

ASF GitHub Bot logged work on BEAM-14383:
-----------------------------------------

                Author: ASF GitHub Bot
            Created on: 10/May/22 16:50
            Start Date: 10/May/22 16:50
    Worklog Time Spent: 10m 
      Work Description: pabloem opened a new pull request, #17601:
URL: https://github.com/apache/beam/pull/17601

   …ws" errors returned by beam.io.WriteToBigQuery"
   
   This reverts commit 358782006e1db86437b3bf61f910db12d654b1e0.
   
   **Please** add a meaningful description for your change here
   
   ------------------------
   
   Thank you for your contribution! Follow this checklist to help us 
incorporate your contribution quickly and easily:
   
    - [ ] [**Choose 
reviewer(s)**](https://beam.apache.org/contribute/#make-your-change) and 
mention them in a comment (`R: @username`).
    - [ ] Format the pull request title like `[BEAM-XXX] Fixes bug in 
ApproximateQuantiles`, where you replace `BEAM-XXX` with the appropriate JIRA 
issue, if applicable. This will automatically link the pull request to the 
issue.
    - [ ] Update `CHANGES.md` with noteworthy changes.
    - [ ] If this contribution is large, please file an Apache [Individual 
Contributor License Agreement](https://www.apache.org/licenses/icla.pdf).
   
   See the [Contributor Guide](https://beam.apache.org/contribute) for more 
tips on [how to make review process 
smoother](https://beam.apache.org/contribute/#make-reviewers-job-easier).
   
   To check the build health, please visit 
[https://github.com/apache/beam/blob/master/.test-infra/BUILD_STATUS.md](https://github.com/apache/beam/blob/master/.test-infra/BUILD_STATUS.md)
   
   GitHub Actions Tests Status (on master branch)
   
------------------------------------------------------------------------------------------------
   [![Build python source distribution and 
wheels](https://github.com/apache/beam/workflows/Build%20python%20source%20distribution%20and%20wheels/badge.svg?branch=master&event=schedule)](https://github.com/apache/beam/actions?query=workflow%3A%22Build+python+source+distribution+and+wheels%22+branch%3Amaster+event%3Aschedule)
   [![Python 
tests](https://github.com/apache/beam/workflows/Python%20tests/badge.svg?branch=master&event=schedule)](https://github.com/apache/beam/actions?query=workflow%3A%22Python+Tests%22+branch%3Amaster+event%3Aschedule)
   [![Java 
tests](https://github.com/apache/beam/workflows/Java%20Tests/badge.svg?branch=master&event=schedule)](https://github.com/apache/beam/actions?query=workflow%3A%22Java+Tests%22+branch%3Amaster+event%3Aschedule)
   
   See [CI.md](https://github.com/apache/beam/blob/master/CI.md) for more 
information about GitHub Actions CI.
   




Issue Time Tracking
-------------------

    Worklog Id:     (was: 768586)
    Time Spent: 4h 50m  (was: 4h 40m)

> Improve "FailedRows" errors returned by beam.io.WriteToBigQuery
> ---------------------------------------------------------------
>
>                 Key: BEAM-14383
>                 URL: https://issues.apache.org/jira/browse/BEAM-14383
>             Project: Beam
>          Issue Type: Improvement
>          Components: io-py-gcp
>            Reporter: Oskar Firlej
>            Priority: P2
>             Fix For: 2.39.0
>
>          Time Spent: 4h 50m
>  Remaining Estimate: 0h
>
> `WriteToBigQuery` pipeline returns `errors` when trying to insert rows that 
> do not match the BigQuery table schema. `errors` is a dictionary that 
> cointains one `FailedRows` key. `FailedRows` is a list of tuples where each 
> tuple has two elements: BigQuery table name and the row that didn't match the 
> schema.
> This can be verified by running the `BigQueryIO deadletter pattern` 
> https://beam.apache.org/documentation/patterns/bigqueryio/
> Using this approach I can print the failed rows in a pipeline. When running 
> the job, logger simultaneously prints out the reason why the rows were 
> invalid. The reason should also be included in the tuple in addition to the 
> BigQuery table and the raw row. This way next pipeline could process both the 
> invalid row and the reason why it is invalid.
> During my reasearch i found a couple of alternate solutions, but i think they 
> are more complex than they need to be. Thats why i explored the beam source 
> code and found the solution to be an easy and simple change.



--
This message was sent by Atlassian Jira
(v8.20.7#820007)

Reply via email to