Hi, I have almost finished the initial packaging of sepp [0]. Beside the sepp program, upstream also provides the tipp program in the same tarball. Basically, tipp classifies sequences using sepp and a collection of alignments and placements data and statistical methods. People installing tipp are invited to download a dataset (approx. 240Mo) [1] which does not belong to the same Github repository and has no license information inside it.
Technically, I guess we might consider creating a sepp-data package with those data, but I also imagine this is not really feasible if we don't have much information about where those data come from, who collected them, ... Based on your experience, would you have some advice on this? My proposal is to let tipp aside and only focus on sepp, which is ready. Thanks a lot, Pierre [0] https://salsa.debian.org/med-team/sepp [1] https://github.com/tandyw/tipp-reference/

