[
https://issues.apache.org/jira/browse/HBASE-7952?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13589187#comment-13589187
]
Raymond Liu commented on HBASE-7952:
------------------------------------
Hi Lars Hofhansl
You are right, beyond micro benchmark, when scan a table, the performance gain
is neglectable.
And actually, there are another line of code in ExplicitColumnTracker did
impact performance a lot, say 30% for real table scan.
as I commented in HBASE-4433, if most of the KV only have one history version.
the include,then seek approaching is much faster than include_and_seek
approaching. Would you mind to take a look upon that issue?
> Remove update() and Improve ExplicitColumnTracker performance.
> --------------------------------------------------------------
>
> Key: HBASE-7952
> URL: https://issues.apache.org/jira/browse/HBASE-7952
> Project: HBase
> Issue Type: Improvement
> Components: regionserver
> Affects Versions: 0.94.1, 0.94.5
> Reporter: Raymond Liu
> Assignee: Raymond Liu
> Fix For: 0.96.0
>
> Attachments: HBASE_7952.patch
>
>
> In ColumnTracker.java, the update() method is not used by anyone now. And no
> one will call checkColumn for different HFiles with update() in between files
> to re-walk through the target columns. All columns will be feed to
> checkColumn() in order.
> So, within ExplicitColumnTracker, the target columns can be optimized to not
> dynamic maintain a changing list of columns yet to match. Instead, just move
> index through it is enough.
> with this optimization to save the time for avoid reconstruct a columns array
> upon each row, the checkColumn method's performance could be improved by
> 10-20%.
--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira