Hi,

I am new to Elasticsearch which I understand can do much more than this... 
but could it be used just for that ?

I am storing 100GB of log files daily. The data scientists require this log 
data to not contain duplicate log lines. Duplicates may come within the 
same log file, with two sequential log files - it's better to expect any 
possible scenario.

What I would like to achieve is to use Elasticsearch to detect & remove the 
duplicate log lines from all logs in an HDFS directory. Can this be done ?

Thank you,
Mihai

-- 
You received this message because you are subscribed to the Google Groups 
"elasticsearch" group.
To unsubscribe from this group and stop receiving emails from it, send an email 
to [email protected].
To view this discussion on the web visit 
https://groups.google.com/d/msgid/elasticsearch/a2dd6d7e-698f-4e03-908f-17358c902f6c%40googlegroups.com.
For more options, visit https://groups.google.com/d/optout.

Reply via email to