rok commented on PR #51680: URL: https://github.com/apache/arrow/pull/51680#issuecomment-6015288122
> > I suppose we could generate stubs for compute kernels with stubgen > > Maybe I am not familiar enough with how this all works, but why do we need stubs for python files? Why not add type annotations in compute.py itself? Or the problem for this file specifically is that it contains dynamically generated functions, and type checkers need to get that information statically? Yes, kernels are never available statically. Can we technically put stub information into .py files? Seems unconventional. > Then we can still only provide a stub file for compute.py, and do inline annotations for the other python files? Yes, and we have two separate ways and locations for storing annotations. I am not against this, but I'd note it complicates things. > > Both dynamic and cython/python stubs will be included into git without docstrings (we agreed on this somewhere in or in vicinity of #32609). We will want to use our [existing docstring populating script](https://github.com/apache/arrow/blob/main/python/scripts/update_stub_docstrings.py) or `docstring-adder` to inject docstrings into stubs at wheel build time. > > Why would we explicitly remove docstrings first (since the stubgen can include them) to then re-inject them with some build time magic? Is that for reducing the size of the git tracked sources? (as long as the files are auto-generated, this might not be too much of a problem?) > > (I checked mentions of "docstring" in #32609, but those are all rather about _wanting_ docstrings) Found the thread https://github.com/apache/arrow/discussions/45919#discussioncomment-14275898 > > cover the dynamically generated surface (kernels) with manual stubs > > I think we could relatively easily also generate those stubs? (we already generate the python code with signature on the fly, then we can also generate annotations on the fly and save that with some extra script to a file) Ok, but we can't have rich types. I don't think we want that. > > What I'm a little worried about is potential namespaces (dynamic, cython and python) overlap which might require extra logic. If this logic gets too complex I would be against using the mix of auto-generated and manual stubs (and slightly prefer manual-only). This would be to keep maintenance cost of stubs as low as possible. > > I think the goal should be "no manual stubs" if we go stubgen(-pyx) If we can somehow annotate dynamically generated symbols then we can have "no manual stubs" otherwise we will always lose rich types. > IMO the problem with this approach is that this first item is the bottleneck. The current manual stubs are super complex, and hard to all get in (the stalled PRs). (unless you would go for starting with lighter manual stubs, and then also gradually enrich those, but lighter manual stubs are essentially the auto-generated ones, potentially) I think the stalled PRs is mostly on me not having full time availability for this. As for the proposed types - do you believe they are overly complex for what's needed downstream? Let's try to set a target type complexity and then find a way to it. From memory types are pretty close to what you need if you want to e.g. infer what will be output type of a function given an input type. If this is something we ultimately want I don't think we can simplify much. If we now throw the internal types away and start with simple types we'll have to recreate internal types later to get to desired functionality which I would like to avoid. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
