This is an automated email from the ASF dual-hosted git repository.
damccorm pushed a commit to branch master
in repository https://gitbox.apache.org/repos/asf/beam.git
The following commit(s) were added to refs/heads/master by this push:
new b4db8d2a11d Fix documentation typos across website, contributor docs,
and Go SDK docs (#40316)
b4db8d2a11d is described below
commit b4db8d2a11d6b2796a98a149ea0c6709d97d6b36
Author: Danny McCormick <[email protected]>
AuthorDate: Mon Sep 28 17:45:44 2026 +0000
Fix documentation typos across website, contributor docs, and Go SDK docs
(#40316)
---
contributor-docs/python-tips.md | 2 +-
contributor-docs/release-guide.md | 8 ++--
sdks/go/BUILD.md | 2 +-
sdks/go/README.md | 2 +-
sdks/go/pkg/beam/runners/prism/README.md | 18 ++++-----
sdks/go/pkg/beam/runners/prism/internal/README.md | 44 +++++++++++-----------
.../www/site/content/en/contribute/dependencies.md | 2 +-
.../dsls/sql/extensions/create-external-table.md | 2 +-
.../en/documentation/io/built-in/snowflake.md | 2 +-
.../content/en/documentation/io/io-standards.md | 4 +-
.../site/content/en/documentation/runners/prism.md | 4 +-
.../content/en/documentation/sdks/typescript.md | 2 +-
.../en/documentation/sdks/yaml-providers.md | 2 +-
.../python/elementwise/enrichment-milvus.md | 2 +-
.../transforms/python/elementwise/mltransform.md | 2 +-
.../python/elementwise/runinference-sklearn.md | 2 +-
.../transforms/python/elementwise/runinference.md | 2 +-
.../en/get-started/mobile-gaming-example.md | 2 +-
18 files changed, 52 insertions(+), 52 deletions(-)
diff --git a/contributor-docs/python-tips.md b/contributor-docs/python-tips.md
index cee96404df3..90e2f283498 100644
--- a/contributor-docs/python-tips.md
+++ b/contributor-docs/python-tips.md
@@ -506,7 +506,7 @@ When we build Python [container images for the Apache Beam
SDK](https://beam.ap
We
[expect](https://github.com/apache/beam/blob/release-2.35.0/sdks/python/container/Dockerfile#L41-L42)
all Beam dependencies (including transitive dependencies, and deps for some of
the 'extra's, like [gcp]) to be specified with exact versions in the
requirements files. When you modify the Python SDK's dependencies in setup.py,
you might need to regenerate the requirements files when or wait until a [PR
updating Python dependency
files](https://github.com/apache/beam/pulls?q=is%3Apr+au [...]
-Regenerate the requirements files by running: `./gradlew
:sdks:python:container:generatePythonRequirementsAll` and commiting the
changes. Execution can take up to 5 min per Python version and is somewhat
resource-demanding. You can also regenerate the dependencies individually per
version with targets like `./gradlew
:sdks:python:container:py38:generatePythonRequirements`.
+Regenerate the requirements files by running: `./gradlew
:sdks:python:container:generatePythonRequirementsAll` and committing the
changes. Execution can take up to 5 min per Python version and is somewhat
resource-demanding. You can also regenerate the dependencies individually per
version with targets like `./gradlew
:sdks:python:container:py38:generatePythonRequirements`.
To run the command successfully, you will need Python interpreters for all
versions supported by Beam. See: [Installing Python
Interpreters](#installing-python-interpreters).
diff --git a/contributor-docs/release-guide.md
b/contributor-docs/release-guide.md
index 7c96465c722..fb30816b2b5 100644
--- a/contributor-docs/release-guide.md
+++ b/contributor-docs/release-guide.md
@@ -779,7 +779,7 @@ as an example.
Use the content of the blog post as the description of the release.
You may now also uncheck the "draft" checkbox.
-This allows it to be visible to non-committers, and makes the assets
publically accessible.
+This allows it to be visible to non-committers, and makes the assets publicly
accessible.
Be sure the release is still marked as a pre-release (not as latest).
@@ -1013,7 +1013,7 @@ If the issue persists, create an infrastructure ticket
for assistance (e.g., htt
Once the tag is uploaded, update the page with the final release tag, and
publish the release notes to Github.
* From the [Beam release page on
Github](https://github.com/apache/beam/releases)
-find and open the release for the final RC tag for for editing.
+find and open the release for the final RC tag for editing.
* Update the release with the final version tag created above.
* Set this version as the latest release, and publish it.
@@ -1351,7 +1351,7 @@ Please review and vote on the release candidate #1 for
the version 2.XX.1. Given
### Revert a commit on a release branch
-The recomended approach is to use `git revert`, for example,
+The recommended approach is to use `git revert`, for example,
```bash
git checkout origin/release-2.62.0
git revert 41215a3116b5e866d1e5b017611a479eeee72df1
@@ -1360,4 +1360,4 @@ git push origin HEAD:release-2.62.0
### How to create a cherry-pick
-More detailes are at
https://cwiki.apache.org/confluence/display/BEAM/Git+Tips#GitTips-Howtocreateacherry-pickpullrequestforanongoingreleasebranch
+More details are at
https://cwiki.apache.org/confluence/display/BEAM/Git+Tips#GitTips-Howtocreateacherry-pickpullrequestforanongoingreleasebranch
diff --git a/sdks/go/BUILD.md b/sdks/go/BUILD.md
index e2606f3597e..bda56254832 100644
--- a/sdks/go/BUILD.md
+++ b/sdks/go/BUILD.md
@@ -33,7 +33,7 @@ In short, the goals are to make both worlds work well.
## Go Modules
-Beam publishes a single Go Module for SDK developement and usage, in the
`sdks` directory.
+Beam publishes a single Go Module for SDK development and usage, in the `sdks`
directory.
This puts all Go code necessary for user pipeline development and for execution
under the same module.
This includes container bootloader code in the Java and Python SDK directories.
diff --git a/sdks/go/README.md b/sdks/go/README.md
index bb7f22934f6..0bf5ae5a902 100644
--- a/sdks/go/README.md
+++ b/sdks/go/README.md
@@ -109,7 +109,7 @@ Crown'd: 1
Note that, when running at Beam HEAD, the Dataflow runner will try to use a
non-existent container `gcr.io/cloud-dataflow/v1beta3/beam_go_sdk:<Beam
version>.dev`.
To address this, you need to push your own SDK harness container image to a
repository (for example, Docker Hub or Google Artifact Registry) and specify
that as the
`<YOUR_SDK_HARNESS_IMAGE_LOCATION>` parameter above.
-For example, running the following from Beam HEAD, will make the container
availble at the location `<repository>/beam_go_sdk`.
+For example, running the following from Beam HEAD, will make the container
available at the location `<repository>/beam_go_sdk`.
```bash
$ ./gradlew :sdks:go:container:docker -Pdocker-repository-root=<repository>
diff --git a/sdks/go/pkg/beam/runners/prism/README.md
b/sdks/go/pkg/beam/runners/prism/README.md
index 0be9ca5617d..08ac3de33ce 100644
--- a/sdks/go/pkg/beam/runners/prism/README.md
+++ b/sdks/go/pkg/beam/runners/prism/README.md
@@ -76,13 +76,13 @@ Here's a non-exhaustive set of variants.
The "default" variant is testing focused, intending to route out issues at
development
time, rather than discovering them on production runners. Notably, this mode
should
-never use fusion, executing each Transform individually and independantly, one
at a time.
+never use fusion, executing each Transform individually and independently, one
at a time.
This variant should be able to execute arbitrary pipelines, correctly, with
clarity and
precision when an error occurs. Other features supported by the SDK should be
enabled by default to
ensure good coverage, such as caches, or RPC reductions like sending elements
in
ProcessBundleRequest and Response, as they should not affect correctness.
Composite
-transforms like Splitable DoFns and Combines should be expanded to ensure
coverage.
+transforms like Splittable DoFns and Combines should be expanded to ensure
coverage.
Additional validations may be added as time goes on.
@@ -96,21 +96,21 @@ executions.
Not Yet Implemented - Illustrative goal.
The "fast" variant is performance focused, intended for local scale execution.
-A psuedo production execution. Fusion optimizations should be performed.
+A pseudo production execution. Fusion optimizations should be performed.
Large PCollection should be offloaded to persistent disk. Bundles should be
dynamically split. Multiple Bundles should be executed simultaneously. And so
on.
Pipelines should execute as swiftly as possible within the bounds of correct
execution.
-### Variant Hightlight: "flink" "dataflow" "spark" AKA Emulations
+### Variant Highlight: "flink" "dataflow" "spark" AKA Emulations
Not Yet Implemented - Illustrative goal.
Emulation variants have the goal of replicating on the local scale,
the behaviors of other runners. Flink execution never "lifts" Combines, and
doesn't dynamically split. Dataflow has different characteristics for batch
-and streaming execution with certain execution charateristics enabled or
+and streaming execution with certain execution characteristics enabled or
disabled.
As Prism is intended to implement all facets of Beam Model execution, the
handlers
@@ -144,12 +144,12 @@ can have features selectively disabled to ensure
* Expands Splittable DoFns
* Process Continuations (AKA Streaming transform support)
* Limited support for Process Continuations
- * Residuals are rescheduled for execution immeadiately.
+ * Residuals are rescheduled for execution immediately.
* The transform must be finite (and eventually return a stop process
continuation)
* Basic Metrics support
* Stand alone execution support
* Web UI available when run as a standalone command.
-* Progess tracking
+* Progress tracking
* Channel Splitting
* Dynamic Splitting
* FnAPI Optimizations
@@ -177,7 +177,7 @@ support users of the Go SDK in testing their pipelines.
Until additional structure is necessary, check the main issue
https://github.com/apache/beam/issues/24789 for the current
status, file an issue for the feature or bug to fix with `[prism]`
-in the title, and refer to the main issue, before begining work
+in the title, and refer to the main issue, before beginning work
to avoid duplication of effort.
If a feature will take a long time, please send a PR to
@@ -189,4 +189,4 @@ Otherwise, ordinary [Beam contribution guidelines
apply](https://beam.apache.org
Once support for containers is implemented, Prism should become a target
for the Java Runner Validation tests, which are the current specification
-for correct runner behavior. This will inform further feature developement.
+for correct runner behavior. This will inform further feature development.
diff --git a/sdks/go/pkg/beam/runners/prism/internal/README.md
b/sdks/go/pkg/beam/runners/prism/internal/README.md
index 684d0a80c51..9ba349ddb94 100644
--- a/sdks/go/pkg/beam/runners/prism/internal/README.md
+++ b/sdks/go/pkg/beam/runners/prism/internal/README.md
@@ -35,7 +35,7 @@ not depend on other parts of the runner. Runner packages can
and do depend on ot
parts of the SDK, such as for Coder handling.
`config` contains configuration parsing and handling. Leaf package.
-Handler configurations are registered by dependant packages.
+Handler configurations are registered by dependent packages.
`urns` contains beam URN strings pulled from the protos. Leaf package.
@@ -136,7 +136,7 @@ WRT necessary restrictions on processing. For example,
stateful stages may
require that only a single inprogress bundle may operate on a given user key
at a time, while aggregations like GroupByKey will only execute when their
windowing strategy dictates, and DoFns with side inputs can only execute when
-all approprate side inputs are ready.
+all appropriate side inputs are ready.
Architecturally, the `ElementManager` is only aware of the properties of the
fused stages, and not their actual relationships with the Beam Protocol
Buffers.
@@ -202,14 +202,14 @@ graph TD;
CheckReady["
For each stage:
see if it has pending elements
- elegible to process with
+ eligible to process with
the current watermark.
"]
Emit["
Output Bundle for Processing
"]
Quiescense{"
- Quiescense check:
+ Quiescence check:
Can the pipeline
make progress?
"}
@@ -269,7 +269,7 @@ So Bundles are produced for a stage based on its watermark
progress, with elemen
A stage's watermarks are determined by it's upstream stages, it's current
watermark state, and it's current
pending elements.
-At the start of a job, all stages are initialized to have watermarks at the
the minimum time, and impulse elements are added to their consuming stages.
+At the start of a job, all stages are initialized to have watermarks at the
minimum time, and impulse elements are added to their consuming stages.
```mermaid
@@ -329,7 +329,7 @@ end
```
-## Bundle Spliting
+## Bundle Splitting
In order to efficiently process data and scale, Beam Runners can use a
combination
of two broad approaches to dividing work, Initial Splitting, and Dynamic
Splitting.
@@ -338,7 +338,7 @@ of two broad approaches to dividing work, Initial
Splitting, and Dynamic Splitti
Initial Splitting is a part of bundle generation, and is decided before bundles
even begin processing. This should take into account the current state
-of the pipeline, oustanding data to be processed, and the current load on the
system.
+of the pipeline, outstanding data to be processed, and the current load on the
system.
Larger bundles require fewer "round trips" to the SDK and batch processing,
but in
general, are processed serially by the runner, leading to higher latency for
downstream
results. Smaller bundles may incur per bundle overhead more frequently, but
can yield lower
@@ -369,7 +369,7 @@ If there hasn't, then a split request is made for half of
the unprocessed work f
The progress interval for the stage (not simply this bundle) is then
increased, to reduce
the frequency of splits if the stage is relatively slow at processing.
The stage's progress interval is decreased if bundles complete so quickly that
no progress requests
-can be made durinng the interval.
+can be made during the interval.
Oversplitting can still occur for this approach, so
https://github.com/apache/beam/issues/32538
proposes incorporating the available execution parallelism into the decision
of whether or not to
@@ -393,9 +393,9 @@ executing a job. This will not include SDK-side threads,
such as those
within containers, or started by an external worker service.
As a rule of thumb, each Bundle is processed on
-an independant goroutine, with a few exceptions.
+an independent goroutine, with a few exceptions.
-jobservices.Server implements a beam JobManagmenent GRPC service. GRPC servers
+jobservices.Server implements a beam JobManagement GRPC service. GRPC servers
have goroutines managed by GRPC itself. Call this G goroutines.
When RunJob is called, the server starts a goroutine with the Job Executor
function.
@@ -430,7 +430,7 @@ The other goroutine blocks until there are no more pending
elements, at which po
cancels the watermark evaluating goroutine, and unblocks the condition
variable, in that order.
Prism's main thread when run as a stand alone command will block forever after
-initializing the beam JobManagment services. Similarly, a built in prism
instance
+initializing the beam JobManagement services. Similarly, a built in prism
instance
in the Go SDK will start a Job Management instance, and the main thread is
blocked
while the job is executing through the "Universal" runner handlers for a
pipeline.
@@ -449,7 +449,7 @@ For each Environment:
* G for the worker's GRPC server.
For each Bundle:
-* 1 to handle bundle execution, up to the configured maxium parallel bundles
(default 8)
+* 1 to handle bundle execution, up to the configured maximum parallel bundles
(default 8)
Letting E be the number of environments in the job, and B maximum number of
parallel bundles:
@@ -459,7 +459,7 @@ Total Goroutines = G + 3 + 9E + E*G + B
Total Goroutines = G(E + 1) + 9E + B + 3
-So for a job J with 1 eviroment, and the default maximum parallel bundles, 8:
+So for a job J with 1 environment, and the default maximum parallel bundles, 8:
Total Goroutines for Job J = G((1) + 1) + 9(1) + (8) + 3
@@ -474,7 +474,7 @@ be the busiest moving data back and forth from the SDK.
A consequence of this approach is the need to take care in locking shared
resources and data
when they may be accessed by multiple goroutines. This, in particular, is all
done in the `ElementManager`
-which has locks for each stage in order to serialze access to it's state. This
state is notably
+which has locks for each stage in order to serialize access to it's state.
This state is notably
accessed by the Bundle goroutines on persisting data back to the
`ElementManager`.
For best performance, we do as much work as possible in the Bundle Goroutines
since they are
@@ -492,7 +492,7 @@ A channel is being used to move ready to execute bundles
from the `ElementManage
This may be unbuffered (the default) which means
serializing how bundles are generated for execution,
and there being at most a single "readyToExecute"
-bundle at a time. An unbufferred channel puts a
+bundle at a time. An unbuffered channel puts a
bottleneck on the job since there may be additional
ready work to execute. On the other hand, it also
allows for bundles to be made larger as more data
@@ -501,7 +501,7 @@ may have arrived.
The channel could be made to be buffered, to allow
multiple bundles to be prepared for execution.
This would lead to lower latency as bundles could be made smaller, and faster
to execute, as it would
-permit pipelineing in work generation, but may lead
+permit pipelining in work generation, but may lead
to higher lock contention and variability in execution.
## Durability Model
@@ -524,19 +524,19 @@ for SDK environments, they'll be on the same machine as
Prism.
* Element: A single value of data to be processed, or a timer to trigger.
* Stage: A fused grouping of one or more transforms with a single parallel
input PCollection,
- zero or more side input PCollecitons, and zero or more output PCollections.
+ zero or more side input PCollections, and zero or more output PCollections.
The engine is unaware of individual user transforms, and relies on the
calling
job executor to configure how stages are related.
* Bundle: An arbitrary non-empty set of elements, to be executed by a stage.
* Upstream Stages: Stages that provide input to the
current stage. Not all stages have upstream stages.
-* Downstream Stages: Stages that depend on input from the current stage. Not
alls tages have downstream stages.
-* Watermark: An event time which relates to the the readiness to process data
in the engine.
+* Downstream Stages: Stages that depend on input from the current stage. Not
all stages have downstream stages.
+* Watermark: An event time which relates to the readiness to process data in
the engine.
Each stage has several watermarks it tracks: Input, Output, and Upstream.
* Upstream Watermark: The minimum output watermark of all stages that
provide input to this stage.
- * Input watermark: The minumum event time of all elements pending or in
progress for this stage.
- * Output watermark: The maxiumum of the current output watermark, the
estimated output watermark (if available), and the minimum of watermark holds.
-* Quiescense: Wether the pipeline is or is able to perform work.
+ * Input watermark: The minimum event time of all elements pending or in
progress for this stage.
+ * Output watermark: The maximum of the current output watermark, the
estimated output watermark (if available), and the minimum of watermark holds.
+* Quiescence: Whether the pipeline is or is able to perform work.
* The pipeline will try to advance all watermarks to infinity, and attempt to
process all pending elements.
* A pipeline will successfully terminate when there are no pending elements
to process,
diff --git a/website/www/site/content/en/contribute/dependencies.md
b/website/www/site/content/en/contribute/dependencies.md
index ac3a669a2d3..367fb4b9997 100644
--- a/website/www/site/content/en/contribute/dependencies.md
+++ b/website/www/site/content/en/contribute/dependencies.md
@@ -74,7 +74,7 @@ For manually identified critical dependency updates, Beam
community members shou
__Dependencies of Java SDK components that may cause issues to other
components if leaked should be vendored.__
-[Vendoring](https://www.ardanlabs.com/blog/2013/10/manage-dependencies-with-godep.html)
is the process of creating copies of third party dependencies. Combined with
repackaging, vendoring allows Beam components to depend on third party
libraries without causing conflicts to other components. Vendoring should be
done in a case-by-case basis since this can increase the total number of
dependencies deployed in user's enviroment.
+[Vendoring](https://www.ardanlabs.com/blog/2013/10/manage-dependencies-with-godep.html)
is the process of creating copies of third party dependencies. Combined with
repackaging, vendoring allows Beam components to depend on third party
libraries without causing conflicts to other components. Vendoring should be
done in a case-by-case basis since this can increase the total number of
dependencies deployed in user's environment.
## Dependency updates and backwards compatibility
diff --git
a/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md
b/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md
index b2ab2f95a4a..1cb45761bfd 100644
---
a/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md
+++
b/website/www/site/content/en/documentation/dsls/sql/extensions/create-external-table.md
@@ -655,7 +655,7 @@ TBLPROPERTIES '{"format: "Excel"}'
* `LOCATION`: The path to the file for Read Mode. The prefix for Write Mode.
* `TBLPROPERTIES`:
* `format`: Optional. Allows you to specify the CSV Format, which
controls
- the field delimeter, quote character, record separator, and other
properties.
+ the field delimiter, quote character, record separator, and other
properties.
See the following table:
<div class="table-container-wrapper">
diff --git a/website/www/site/content/en/documentation/io/built-in/snowflake.md
b/website/www/site/content/en/documentation/io/built-in/snowflake.md
index 6ca61781c60..eab82f194b4 100644
--- a/website/www/site/content/en/documentation/io/built-in/snowflake.md
+++ b/website/www/site/content/en/documentation/io/built-in/snowflake.md
@@ -571,7 +571,7 @@ public static SnowflakeIO.UserDataMapper<Long>
getCsvMapper() {
{{< /highlight >}}
### Additional write options
#### Transformation query
-The `.withQueryTransformation()` option for the `write()` operation accepts a
SQL query as a String value, which will be performed while transfering data
staged in CSV files directly to the target Snowflake table. For information
about the transformation SQL syntax, see the [Snowflake
Documentation](https://docs.snowflake.net/manuals/sql-reference/sql/copy-into-table.html#transformation-parameters).
+The `.withQueryTransformation()` option for the `write()` operation accepts a
SQL query as a String value, which will be performed while transferring data
staged in CSV files directly to the target Snowflake table. For information
about the transformation SQL syntax, see the [Snowflake
Documentation](https://docs.snowflake.net/manuals/sql-reference/sql/copy-into-table.html#transformation-parameters).
Usage:
{{< highlight >}}
diff --git a/website/www/site/content/en/documentation/io/io-standards.md
b/website/www/site/content/en/documentation/io/io-standards.md
index 1f1ee505b0b..204275fb5d2 100644
--- a/website/www/site/content/en/documentation/io/io-standards.md
+++ b/website/www/site/content/en/documentation/io/io-standards.md
@@ -296,7 +296,7 @@ The I/O Connector development guidelines are written with
the following principl
</td>
<td>
<p>An I/O should rarely rely on a PipelineOptions subclass to tune
internal parameters.
- <p>If neccesary, a connector-related pipeline options class should:
+ <p>If necessary, a connector-related pipeline options class should:
<ul>
<li>Document clearly, for each option, the effect it has and why
one may modify it.
<li>Option names must be namespaced to avoid collisions
@@ -1296,7 +1296,7 @@ When possible, unit tests are favored over integration
tests due to faster execu
<p>Sink batching test
</td>
<td>
- <p>Make sure that sinks batch data before writing if the sinks
performace batching for performance reasons.
+ <p>Make sure that sinks batch data before writing if the sinks
perform batching for performance reasons.
</td>
<td>
<p><a
href="https://github.com/apache/beam/blob/c57c983c8ae7d84926f9cf42f7c40af8eaf60545/sdks/java/io/google-cloud-platform/src/test/java/org/apache/beam/sdk/io/gcp/spanner/SpannerIOWriteTest.java#L1200">SpannerIOWriteTest.testBatchFn_cells</a>
diff --git a/website/www/site/content/en/documentation/runners/prism.md
b/website/www/site/content/en/documentation/runners/prism.md
index a898a691812..008124fd2dd 100644
--- a/website/www/site/content/en/documentation/runners/prism.md
+++ b/website/www/site/content/en/documentation/runners/prism.md
@@ -131,14 +131,14 @@ Simply unzip, and execute.
This approach requires a [recent version of Go installed](https://go.dev/dl/).
This is recommended if you only want to run Prism on your local machine.
-You can insall Prism with `go install`:
+You can install Prism with `go install`:
```sh
go install github.com/apache/beam/sdks/v2/go/cmd/prism@latest
prism
```
-Or simply build and execute the binary immeadiately using `go run`:
+Or simply build and execute the binary immediately using `go run`:
```sh
go run github.com/apache/beam/sdks/v2/go/cmd/prism@latest
diff --git a/website/www/site/content/en/documentation/sdks/typescript.md
b/website/www/site/content/en/documentation/sdks/typescript.md
index 1f201f0de0f..69c52c167ee 100644
--- a/website/www/site/content/en/documentation/sdks/typescript.md
+++ b/website/www/site/content/en/documentation/sdks/typescript.md
@@ -87,7 +87,7 @@ themselves, producing multiple outputs is done by following
with a new
`PCollection<{a?: AType, b: BType, ... }>` and produces an object
`{a: PCollection<AType>, b: PCollection<BType>, ...}`.
-* JavaScript supports (and encourages) an asynchronous programing model, with
+* JavaScript supports (and encourages) an asynchronous programming model, with
many libraries requiring use of the async/await paradigm.
As there is no way (by design) to go from the asynchronous style back to
the synchronous style, this needs to be taken into account
diff --git a/website/www/site/content/en/documentation/sdks/yaml-providers.md
b/website/www/site/content/en/documentation/sdks/yaml-providers.md
index 51d4634f0e1..743f7f38ceb 100644
--- a/website/www/site/content/en/documentation/sdks/yaml-providers.md
+++ b/website/www/site/content/en/documentation/sdks/yaml-providers.md
@@ -237,7 +237,7 @@ in the same format as those inlined in this providers block.
See, for example, the provider listing [here](
https://github.com/apache/beam-starter-python-provider/blob/main/examples/provider_listing.yaml).
-In fact, this is how many of the the built in transforms are declared,
+In fact, this is how many of the built in transforms are declared,
see for example the [builtin io listing file](
https://github.com/apache/beam/blob/master/sdks/python/apache_beam/yaml/standard_io.yaml).
diff --git
a/website/www/site/content/en/documentation/transforms/python/elementwise/enrichment-milvus.md
b/website/www/site/content/en/documentation/transforms/python/elementwise/enrichment-milvus.md
index f57c2b627ec..045e94e5417 100644
---
a/website/www/site/content/en/documentation/transforms/python/elementwise/enrichment-milvus.md
+++
b/website/www/site/content/en/documentation/transforms/python/elementwise/enrichment-milvus.md
@@ -54,7 +54,7 @@ Output:
{{< code_sample
"sdks/python/apache_beam/examples/snippets/transforms/elementwise/enrichment_test.py"
enrichment_with_milvus >}}
{{< /highlight >}}
-## Notebook exmaple
+## Notebook example
<a
href="https://colab.research.google.com/github/apache/beam/blob/master/examples/notebooks/beam-ml/milvus_enrichment_transform.ipynb"
target="_blank">
<img src="https://colab.research.google.com/assets/colab-badge.svg"
alt="Open In Colab" width="150" height="auto" style="max-width: 100%"/>
diff --git
a/website/www/site/content/en/documentation/transforms/python/elementwise/mltransform.md
b/website/www/site/content/en/documentation/transforms/python/elementwise/mltransform.md
index 2eaecbd5a9b..f7758869e4d 100644
---
a/website/www/site/content/en/documentation/transforms/python/elementwise/mltransform.md
+++
b/website/www/site/content/en/documentation/transforms/python/elementwise/mltransform.md
@@ -55,7 +55,7 @@ MLTransform(transforms=transforms,
write_artifact_location=write_artifact_locati
The transforms passed to `MLTransform` are applied sequentially on the
dataset. `MLTransform` expects a dictionary and returns a transformed row
object with NumPy arrays.
## Examples
-The following examples demonstrate how to to create pipelines that use
`MLTransform` to preprocess data.
+The following examples demonstrate how to create pipelines that use
`MLTransform` to preprocess data.
`MLTransform` can do a full pass on the dataset, which is useful when you need
to transform a single element only after analyzing the entire dataset.
The first two examples require a full pass over the dataset to complete the
data transformation.
diff --git
a/website/www/site/content/en/documentation/transforms/python/elementwise/runinference-sklearn.md
b/website/www/site/content/en/documentation/transforms/python/elementwise/runinference-sklearn.md
index af0ce9bd931..e7b7fd46f7a 100644
---
a/website/www/site/content/en/documentation/transforms/python/elementwise/runinference-sklearn.md
+++
b/website/www/site/content/en/documentation/transforms/python/elementwise/runinference-sklearn.md
@@ -27,7 +27,7 @@ limitations under the License.
</tr>
</table>
-The following examples demonstrate how to to create pipelines that use the
Beam RunInference API and Sklearn.
+The following examples demonstrate how to create pipelines that use the Beam
RunInference API and Sklearn.
## Example 1: Sklearn unkeyed model
diff --git
a/website/www/site/content/en/documentation/transforms/python/elementwise/runinference.md
b/website/www/site/content/en/documentation/transforms/python/elementwise/runinference.md
index 0f3cacf1d74..1759a07516f 100644
---
a/website/www/site/content/en/documentation/transforms/python/elementwise/runinference.md
+++
b/website/www/site/content/en/documentation/transforms/python/elementwise/runinference.md
@@ -29,7 +29,7 @@ limitations under the License.
</tr>
</table>
-Uses models to do local and remote inference. A `RunInference` transform
performs inference on a `PCollection` of examples using a machine learning (ML)
model. The transform outputs a `PCollection` that contains the input examples
and output predictions. Avaliable in Apache Beam 2.40.0 and later versions.
+Uses models to do local and remote inference. A `RunInference` transform
performs inference on a `PCollection` of examples using a machine learning (ML)
model. The transform outputs a `PCollection` that contains the input examples
and output predictions. Available in Apache Beam 2.40.0 and later versions.
For more information about Beam RunInference APIs, see the [About Beam
ML](https://beam.apache.org/documentation/ml/about-ml) page and the
[RunInference API
pipeline](https://github.com/apache/beam/tree/master/sdks/python/apache_beam/examples/inference)
examples.
diff --git a/website/www/site/content/en/get-started/mobile-gaming-example.md
b/website/www/site/content/en/get-started/mobile-gaming-example.md
index 63be47688be..4b46697a590 100644
--- a/website/www/site/content/en/get-started/mobile-gaming-example.md
+++ b/website/www/site/content/en/get-started/mobile-gaming-example.md
@@ -60,7 +60,7 @@ Because some of our example pipelines use data files (like
logs from the game se
For pipelines that read unbounded game data from an unbounded source, the data
source sets the intrinsic
[timestamp](/documentation/programming-guide/#element-timestamps) for each
PCollection element to the appropriate event time.
-The Mobile Gaming example pipelines vary in complexity, from simple batch
analysis to more complex pipelines that can perform real-time analysis and
abuse detection. This section walks you through each example and demonstrates
how to use Beam features like windowing and triggers to expand your pipeline's
capabilites.
+The Mobile Gaming example pipelines vary in complexity, from simple batch
analysis to more complex pipelines that can perform real-time analysis and
abuse detection. This section walks you through each example and demonstrates
how to use Beam features like windowing and triggers to expand your pipeline's
capabilities.
## UserScore: Basic Score Processing in Batch