taiyang-li commented on code in PR #2048:
URL: https://github.com/apache/orc/pull/2048#discussion_r1815884042
##########
c++/include/orc/Reader.hh:
##########
@@ -605,6 +612,26 @@ namespace orc {
*/
virtual std::map<uint32_t, BloomFilterIndex> getBloomFilters(
uint32_t stripeIndex, const std::set<uint32_t>& included) const = 0;
+
+ /**
+ * Get the input stream for the ORC file.
+ */
+ virtual InputStream* getStream() const = 0;
+
+ /**
+ * Get the footer of the ORC file.
+ */
+ virtual const proto::Footer* getFooter() const = 0;
+
+ /**
+ * Get the schema of the ORC file.
+ */
+ virtual const proto::Metadata* getMetadata() const = 0;
+
+ virtual void preBuffer(const std::vector<int>& stripes, const
std::list<uint64_t>& includeTypes,
Review Comment:
@ffacs it is a typical use case in spark. ORC files may be very large, and
spark will split it into several file splits, each file split contains several
stripes, we can even discard some candidate stripes by stripe statistics after
that.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]