Hi Wes et al,

I did a bit of thinking on this, speciffically the schema design, and wrote
up a draft [1].
Further, I met Ralf (cced) at EuroPython and it turns out he's faced
similar issues, so I proposed we pool resources.
We started a weekly sync and a temporary repo [2] to collect user stories
and requirements before starting the actual implementation. (If anyone is
interested please let me know and I'll send you an invite.)

I expect that once we have an approximately complete agreement on
requirements, we will proceed to implement components. If possible I'd like
to start with conbench's storage layer.

> I’m working on but I’m at a point where I want to start making sweeping
changes to the UI and I don’t want to be disruptive if folks form Arrow
want to use it

This sounds great and we definitely want innovation and better ergonomics!

[1]
https://docs.google.com/document/d/1ki5Y3b-Fqd9XZpKYBErbmV_Ups4jrwgr7B40W76kTKY/edit?usp=sharing
[2] https://github.com/rok/benchx

Rok

On Wed, Aug 26, 2026 at 10:07 PM Wes McKinney <[email protected]> wrote:

> Hi all,
>
> I wanted to see if there is any interested in engaging on this? If not, I
> can just go my own way and call the new codebase something other than
> Conbench (I am using it to benchmark agentsview.io and some other projects
> I’m working on but I’m at a point where I want to start making sweeping
> changes to the UI and I don’t want to be disruptive if folks form Arrow
> want to use it). I’d prefer to move toward a model where much of the data
> lives in Parquet files queried by DuckDB for example rather than Postgres
> being the system of record.
>
> We are in a weird time at the moment since software is so easy to make if
> you have a clear idea of what you want, but it requires human engagement:
> the agents won’t make anything good without an dialed-in operator in the
> loop.
>
> Thanks
> Wes
>
> On Wed, Jul 8, 2026 at 14:22 Rok Mihevc <[email protected]> wrote:
>
> > We're seeing old [1] and new [2] conbench time out. I'm looking into
> > backend optimization to see if we can resolve our main pain point before
> > proceeding.
> >
> > [1] https://conbench.arrow-dev.org/
> > [2] https://conbench-v2.arrow-dev.org/
> >
> > Rok
> >
> > On Wed, Jul 1, 2026 at 11:20 AM Antoine Pitrou <[email protected]>
> wrote:
> >
> > >
> > > Important things for me in conbench:
> > >
> > > * regression detection, which is quite solid (the algorithm has been
> > > tweaked a lot AFAIK):
> > > https://github.com/apache/arrow/runs/84305379739
> > >
> > > * benchmark result pages and especially the "compare with baseline run
> > > from fork point commit", e.g.
> > > https://conbench.arrow-dev.org/runs/457fe991307a42d786798829edbc29f9/
> > >
> > > * the comparison pages such as
> > >
> > >
> >
> https://conbench.arrow-dev.org/compare/runs/1b61ec2a2670462f8c788da25fda3ca6...4ab7d6720b8c4e698e643f35819745c7/
> > > ; currently they are quite slow to compute and often time out
> > >
> > > Note some UI improvements would be warranted (especially: better
> > > filtering and/or display for benchmark results), we can discuss that
> too.
> > >
> > > Regards
> > >
> > > Antoine.
> > >
> > >
> > > Le 27/06/2026 à 03:24, Wes McKinney a écrit :
> > > > I can update the chart to be more in line with the old one.
> > > >
> > > > I haven't done a great deal of work to enhance the UI (a lot of this
> > > > development was unattended and based on a mandate to rebuild the
> > > > backend in Go and rebuild the frontend on modern web technology, but
> > > > not make substantial changes). Using modern web technology (vite /
> > > > svelte) means it is much easier to add new things and make them feel
> > > > polished and nice, so the question is really what would be useful to
> > > > have?
> > > >
> > > > On Fri, Jun 26, 2026 at 7:56 PM Rok Mihevc <[email protected]>
> > wrote:
> > > >>
> > > >> It's night here so I appreciate the new darker palette.
> > > >> As for result readability I prefer the old interactive graph because
> > it
> > > has
> > > >> classified points, trendline and bands.
> > > >>
> > > >> old:
> > > >>
> > >
> >
> https://conbench.arrow-dev.org/benchmark-results/06a3edbc2d1f7dc48000633dc612769e/
> > > >> new::
> > > >>
> > >
> >
> https://conbench-v2.arrow-dev.org/benchmarks/history/06a3edbc2d1f7dc48000633dc612769e
> > > >>
> > > >> Any new feature that you'd point out for improved ergonomics?
> > > >>
> > > >> Rok
> > > >>
> > > >>
> > > >> On Sat, Jun 27, 2026 at 2:22 AM Wes McKinney <[email protected]>
> > > wrote:
> > > >>
> > > >>> TL;DR take a look at https://conbench-v2.arrow-dev.org.
> > > >>>
> > > >>> This was written by Codex to summarize the state of things:
> > > >>>
> > > >>> Hi all,
> > > >>>
> > > >>> I’ve pushed the Conbench v2 work for review in the following
> places:
> > > >>>
> > > >>> Conbench v2 app/server/docs:
> > > >>> https://github.com/conbench/conbench/tree/experimental-v2
> > > >>>
> > > >>> Buildkite, Terraform, and CI adapter work:
> > > >>>
> > https://github.com/wesm/arrow-benchmarks-ci/tree/v2-conbench-ci-report
> > > >>>
> > > >>> Python benchmark payload work:
> > > >>> https://github.com/wesm/benchmarks/tree/v2-conbench-submit
> > > >>>
> > > >>> R benchmark payload work:
> > > >>> https://github.com/wesm/arrowbench/tree/v2-conbench-payloads
> > > >>>
> > > >>> There is also a live read-only evaluator running here:
> > > >>>
> > > >>> https://conbench-v2.arrow-dev.org
> > > >>>
> > > >>> Temporary docs are published here:
> > > >>>
> > > >>> https://wesm.github.io/conbench-tmp/
> > > >>>
> > > >>> The goal of this work is to provide a concrete migration path for
> > > >>> Conbench v2 while preserving the existing production database
> schema.
> > > >>> The evaluator is intended for review and experimentation only; it
> > > >>> should not modify the production database.
> > > >>>
> > > >>> Please take a look at the app, docs, and workflow changes. The main
> > > >>> things to review are whether the new UI is useful for Arrow
> > > >>> maintainers, whether the benchmark reporting path is
> understandable,
> > > >>> and whether the migration approach looks practical for the existing
> > > >>> Arrow benchmarking workflows.
> > > >>>
> > > >>> On Tue, Jun 23, 2026 at 5:34 PM Wes McKinney <[email protected]>
> > > wrote:
> > > >>>>
> > > >>>> If you can give me access to the servers and cloud resources in
> > > question
> > > >>> (ie the machines that currently run the benchmarks), I can
> implement
> > > the
> > > >>> necessary code adaptations and test things working end to end, and
> > > stand up
> > > >>> a parallel deployment of the application against RDS to enable
> better
> > > >>> evaluation of the UI ergonomics. Perhaps we can coordinate offline
> > and
> > > come
> > > >>> back to the community with a report once the implementation is
> closer
> > > to
> > > >>> being able to “throw the switch”.
> > > >>>>
> > > >>>> On Tue, Jun 23, 2026 at 17:02 Rok Mihevc <[email protected]>
> > > wrote:
> > > >>>>>
> > > >>>>> Current architecture is not optimal or modern, but its
> maintenance
> > > cost
> > > >>>>> is well known and currently manageable. But I'm happy with
> changes
> > > that
> > > >>>>> move us to a more maintainable state and am willing to assist
> with
> > > the
> > > >>>>> transition.
> > > >>>>>
> > > >>>>> On a personal note, I'm hesitant to sign up for maintaining a
> > > deployment
> > > >>>>> of software that's not built yet.
> > > >>>>> So I'm just curious about who "owns" Arrow's conbench deployment
> > > >>>>> if I cannot commit to it.
> > > >>>>>
> > > >>>>> Rok
> > > >>>>>
> > > >>>>> On Tue, Jun 23, 2026 at 4:40 PM Wes McKinney <
> [email protected]>
> > > >>> wrote:
> > > >>>>>
> > > >>>>>> I think the objective is to create a more modern foundation
> while
> > > also
> > > >>>>>> improving performance. I think this looks like:
> > > >>>>>>
> > > >>>>>> * faster, easier to develop and deploy backup (use Go or Rust to
> > > >>> create
> > > >>>>>> static binaries: Go is good for backend services like this)
> > > >>>>>> * use modern web technologies versus generating pages with Jinja
> > > >>> templates
> > > >>>>>> * Use things like DuckDB and Parquet to scale result storage
> while
> > > >>>>>> improving performance
> > > >>>>>> * Add many more UI features to make the results most useful to
> > > >>> maintainers
> > > >>>>>>
> > > >>>>>> Like I said, I’m happy to fulfill feature requests and
> contribute
> > > >>>>>> development with agents, if it isn’t interesting I’m also fine
> to
> > go
> > > >>> my own
> > > >>>>>> way.
> > > >>>>>>
> > > >>>>>> I think Buildkite is fine for job scheduling and management for
> > > now, I
> > > >>>>>> don’t think this system currently wants to own a task queue /
> > > durable
> > > >>>>>> execution state for workers, though it could grow this
> capability
> > in
> > > >>> the
> > > >>>>>> future (workers would have to poll the server for jobs to take).
> > > >>>>>>
> > > >>>>>> On Tue, Jun 23, 2026 at 09:07 Antoine Pitrou <
> [email protected]>
> > > >>> wrote:
> > > >>>>>>
> > > >>>>>>>
> > > >>>>>>> Hi,
> > > >>>>>>>
> > > >>>>>>> I think the main question here is: what are we trying to do?
> > > >>>>>>>
> > > >>>>>>> Currently, the main operational issue with conbench is the
> > slowness
> > > >>> of
> > > >>>>>>> the web UI, due to the large database size and that it's not
> > > >>> normalized
> > > >>>>>>> (some queries take much longer than they should).
> > > >>>>>>>
> > > >>>>>>> I can't speak about the maintenance / reliability aspects,
> > though.
> > > >>>>>>>
> > > >>>>>>> Regards
> > > >>>>>>>
> > > >>>>>>> Antoine.
> > > >>>>>>>
> > > >>>>>>>
> > > >>>>>>> Le 04/06/2026 à 16:19, Wes McKinney a écrit :
> > > >>>>>>>> hi all,
> > > >>>>>>>>
> > > >>>>>>>> I saw that conbench.ursa.dev has been down and I had a need
> to
> > > >>> set up
> > > >>>>>>>> some continuous project benchmarks, and was interested in
> doing
> > > >>>>>>>> development on Conbench (well, having my agents do development
> > on
> > > >>>>>>>> Conbench), and was interested in the following:
> > > >>>>>>>>
> > > >>>>>>>> 1) is there interest in migrating the historical Arrow
> conbench
> > > >>> data
> > > >>>>>>>> to a new server, has that been preserved somewhere? I'll
> > probably
> > > >>>>>>>> rewrite the conbench backend in Go and give it a client CLI
> for
> > > >>>>>>>> submitting new data or querying old data.
> > > >>>>>>>>
> > > >>>>>>>> 2) are there other users of conbench (conbench/conbench) that
> > > >>> anyone
> > > >>>>>>>> is aware of? I'd be done doing in-situ development in that
> > > >>> repository
> > > >>>>>>>> or setting up a conbench-v2 project.
> > > >>>>>>>>
> > > >>>>>>>> No particular urgency but if anyone has opinions let me know!
> > > >>>>>>>>
> > > >>>>>>>> thanks,
> > > >>>>>>>> Wes
> > > >>>>>>>
> > > >>>>>>>
> > > >>>>>>
> > > >>>
> > >
> > >
> >
>

Reply via email to