Good stuff. I see the patterns you are discussing on those slides as coming up a lot for us.
On 12/9/15 6:25 PM, "Sheetal Dolas" <[email protected]> wrote: >Additionally this might give some ideas (slide 14 onwards) for handling >these types of problems. > >http://www.slideshare.net/Hadoop_Summit/design-patterns-for-real-time-stre >aming-data-analytics > > > > > > >On 12/9/15, 4:20 PM, "James Sirota" <[email protected]> wrote: > >>So the nature of the problem was that as we were processing ~1.3 million >>of messages per second the time it took for the in-memory DB to update >>caused the Storm tuples to back up to a point where this would bring >>down the topology. We also had problems during initialization. I don’t >>know if this feature exists now, but at the time we couldn’t figure out >>a way to have the topology deploy and wait for all the instances of Geo >>bolt to finish reading their data and signal back that they were ready. >>So at initialization they would get blasted with tuples and fall over. >>We solved that problem at the time by delaying our ingest 30 seconds to >>give the topology a chance to fully come up. But eventually we decided >>we needed to simplify things so we abandoned the in-memory route. >> >>Thanks, >>James >> >> >> >>On 12/9/15, 5:35 PM, "Bryan Taylor" <[email protected]> wrote: >> >>> >>>The GeoLite site says they update once a month, so I assume something >>>can >>>check for this and grab the new file. It seems like a fun problem to >>>have >>>this also trigger a rebuild of the in-memory cache and swap it out live. >>>This seems like it would be a useful streaming enrichment pattern, where >>>the configuration data for the enrichment changes. >>> >>>This does raise another interesting question about what do we expect the >>>memory profile of the steam processing to be. 70Mb + 40Mb isn't that big >>>by itself, but when is it worth it? and how do operators take advantage >>>of >>>more system memory if they have it. >>> >>> >>>On 12/9/15 4:32 PM, "James Sirota" <[email protected]> wrote: >>> >>>>Hi Bryan, >>>> >>>>We had HSQLDB at one point, but we were struggling to make these bolts >>>>reliable. Also, the geo data needs to be periodically updated and it¹s >>>>easier to do when it¹s decoupled. >>>> >>>>Thanks, >>>>James >>>> >>>> >>>> >>>>On 12/9/15, 4:26 PM, "Bryan Taylor" <[email protected]> wrote: >>>> >>>>>[Sorry about the stupid COMMERCIAL: tag being added - I'm trying to >>>>>fix >>>>>with our email folks] >>>>> >>>>>Nice. >>>>> >>>>>I was just looking around at the geotagging enrichment adapter. The >>>>>city >>>>>data is split between the two files of 70Mb and 40Mb sizes. It seems >>>>>like >>>>>the data is small enough to just load it all into memory. This would >>>>>eliminate two SQL queries for every event. >>>>> >>>>>Bryan >>>>> >>>>>On 12/9/15 2:14 PM, "Mark Bittmann" <[email protected]> wrote: >>>>> >>>>>>Hi Bryan, >>>>>> >>>>>>For automation, B23 is planning to contribute Ansible scripts to >>>>>>deploy >>>>>>the Metron stack. The playbooks use Ambari blueprints for the hadoop >>>>>>ecosystem. We also install Elasticsearch, configure the legacy >>>>>>OpenSOC >>>>>>UI >>>>>>(based on Kibana/nodejs), create directories in hdfs, populate a >>>>>>MySQL >>>>>>database for geotagging. We template the OpenSOC_Config files so >>>>>>that we >>>>>>can use variable injection for the different services: Zookeeper, >>>>>>Hbase, >>>>>>Elasticsearch, MySQL, etc. Everyone's deployment might be slightly >>>>>>different, but I think this will be a really good start. >>>>>> >>>>>>I'm in the process of decoupling the scripts from our internal >>>>>>tooling - >>>>>>I should be able to make available the Ambari stuff later this week. >>>>>>Once >>>>>>we merge the disparate forks of the Cisco codebase, there will be >>>>>>some >>>>>>work to bring the playbooks up to date (i.e., use of Storm Flux), but >>>>>>not >>>>>>a ton. >>>>>> >>>>>>Mark >>>>>> >>>>>> >>>>>> >>>>>>On 12/9/15, 2:48 PM, "Bryan Taylor" <[email protected]> wrote: >>>>>> >>>>>>> >>>>>>>Seems like a good list. I'm probably not your UI guy, but I'll be >>>>>>>happy >>>>>>>to >>>>>>>help with any of the others. Test coverage is probably a good place >>>>>>>for >>>>>>>me >>>>>>>to start learning my way around. Do we have CI set up? I see a >>>>>>>travis >>>>>>>file >>>>>>>in the code. >>>>>>> >>>>>>>Regarding automation deployments, are we targeting Ambari or >>>>>>>something >>>>>>>else? >>>>>>> >>>>>>>On the hadoop component compatibility, I see from the >>>>>>>opens-streaming >>>>>>>pom >>>>>>>that we are using Storm-0.9.2, Kafka 0.8.0, Hadoop 2.2.0, and HBase >>>>>>>0.98.0-hadoop2. These are all several iterations old. How >>>>>>>aggressive do >>>>>>>we >>>>>>>want to be, generally, with tracking new releases? Are Hive, Flume, >>>>>>>and >>>>>>>Spark also going to up rev for us? >>>>>>> >>>>>>>Bryan >>>>>>> >>>>>>>On 12/9/15 12:22 PM, "James Sirota" <[email protected]> wrote: >>>>>>> >>>>>>>>Hi Brian, >>>>>>>> >>>>>>>>Welcome. Glad to have you contribute. There will be changes to >>>>>>>>the >>>>>>>>code >>>>>>>>base that the community will contribute back shortly. We are >>>>>>>>waiting >>>>>>>>for >>>>>>>>the Jira to be setup so the backlog can be created and voted on. I >>>>>>>>think >>>>>>>>the overall feeling is that we need to make the code base >>>>>>>>compatible >>>>>>>>with >>>>>>>>the latest version of HDP, automate deployments, increase test >>>>>>>>coverage, >>>>>>>>and start working on a new UI. There may be more significant >>>>>>>>architectural changes to the code base, but we need to get the >>>>>>>>essential >>>>>>>>items knocked out before we go there. >>>>>>>> >>>>>>>>Thanks, >>>>>>>>James >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>> >>>>>>>>On 12/9/15, 11:52 AM, "Bryan Taylor" <[email protected]> wrote: >>>>>>>> >>>>>>>>>Hi Folks, >>>>>>>>> >>>>>>>>>I just joined the list and thought I'd say "hi". I work at >>>>>>>>>Rackspace >>>>>>>>>and >>>>>>>>>will be joining Andrew Hartnett's team there and hacking on >>>>>>>>>metron. >>>>>>>>>This >>>>>>>>>is my first ASF project and I'm looking forward to being part of >>>>>>>>>this >>>>>>>>>community. >>>>>>>>> >>>>>>>>>I'm curious what the development vision is for metron. What do >>>>>>>>>people >>>>>>>>>like and not like about the codebase? I gather this code >>>>>>>>>transitioned >>>>>>>>>from a Cisco internal project and is incubating now at the ASF. >>>>>>>>>Are >>>>>>>>>there any code changes that need to be made to support incubation? >>>>>>>>> >>>>>>>>>Bryan >>>>>>> >>>>> >>>>> >>> >>>
