Hi, On Fri, Apr 27, 2012 at 1:19 PM, Roland Scheidegger <[email protected]> wrote: > This adds a hand-optimized assembly version for get_cabac much like the > existing one, but it works if the table offsets are RIP-relative. > Compared to the non-RIP-relative version this adds 2 lea instructions > and it needs one extra register. > There is a surprisingly large performance improvement over the c version (more > so than the generated assembly seems to suggest) just in get_cabac, I measured > roughly 40% faster for get_cabac on a K8. However, overall the difference is > not that big, I measured roughly 5% on a test clip on a K8 and a Core2. > Hopefully it still compiles on x86 32bit... > Now that only one table is used, there's some chance even darwin as compiles > this (apparently the label arithmetic used previously doesn't work if it > involves symbols defined in a different file, thanks to Ronald S. Bultje for > helping me with this). > --- > libavcodec/h264_cabac.c | 2 +- > libavcodec/x86/cabac.h | 90 ++++++++++++++++++++++++++++++++++++++++--- > libavcodec/x86/h264_i386.h | 53 ++++++++++++++++++-------- > 3 files changed, 121 insertions(+), 24 deletions(-)
Speedup confirmed (indeed 5%, impressive!), very nice work overall. Pushed, and thank you! Ronald _______________________________________________ libav-devel mailing list [email protected] https://lists.libav.org/mailman/listinfo/libav-devel
