Additionally this might give some ideas (slide 14 onwards) for handling these types of problems.
http://www.slideshare.net/Hadoop_Summit/design-patterns-for-real-time-streaming-data-analytics On 12/9/15, 4:20 PM, "James Sirota" <[email protected]> wrote: >So the nature of the problem was that as we were processing ~1.3 million of >messages per second the time it took for the in-memory DB to update caused the >Storm tuples to back up to a point where this would bring down the topology. >We also had problems during initialization. I don’t know if this feature >exists now, but at the time we couldn’t figure out a way to have the topology >deploy and wait for all the instances of Geo bolt to finish reading their data >and signal back that they were ready. So at initialization they would get >blasted with tuples and fall over. We solved that problem at the time by >delaying our ingest 30 seconds to give the topology a chance to fully come up. > But eventually we decided we needed to simplify things so we abandoned the >in-memory route. > >Thanks, >James > > > >On 12/9/15, 5:35 PM, "Bryan Taylor" <[email protected]> wrote: > >> >>The GeoLite site says they update once a month, so I assume something can >>check for this and grab the new file. It seems like a fun problem to have >>this also trigger a rebuild of the in-memory cache and swap it out live. >>This seems like it would be a useful streaming enrichment pattern, where >>the configuration data for the enrichment changes. >> >>This does raise another interesting question about what do we expect the >>memory profile of the steam processing to be. 70Mb + 40Mb isn't that big >>by itself, but when is it worth it? and how do operators take advantage of >>more system memory if they have it. >> >> >>On 12/9/15 4:32 PM, "James Sirota" <[email protected]> wrote: >> >>>Hi Bryan, >>> >>>We had HSQLDB at one point, but we were struggling to make these bolts >>>reliable. Also, the geo data needs to be periodically updated and it¹s >>>easier to do when it¹s decoupled. >>> >>>Thanks, >>>James >>> >>> >>> >>>On 12/9/15, 4:26 PM, "Bryan Taylor" <[email protected]> wrote: >>> >>>>[Sorry about the stupid COMMERCIAL: tag being added - I'm trying to fix >>>>with our email folks] >>>> >>>>Nice. >>>> >>>>I was just looking around at the geotagging enrichment adapter. The city >>>>data is split between the two files of 70Mb and 40Mb sizes. It seems like >>>>the data is small enough to just load it all into memory. This would >>>>eliminate two SQL queries for every event. >>>> >>>>Bryan >>>> >>>>On 12/9/15 2:14 PM, "Mark Bittmann" <[email protected]> wrote: >>>> >>>>>Hi Bryan, >>>>> >>>>>For automation, B23 is planning to contribute Ansible scripts to deploy >>>>>the Metron stack. The playbooks use Ambari blueprints for the hadoop >>>>>ecosystem. We also install Elasticsearch, configure the legacy OpenSOC >>>>>UI >>>>>(based on Kibana/nodejs), create directories in hdfs, populate a MySQL >>>>>database for geotagging. We template the OpenSOC_Config files so that we >>>>>can use variable injection for the different services: Zookeeper, Hbase, >>>>>Elasticsearch, MySQL, etc. Everyone's deployment might be slightly >>>>>different, but I think this will be a really good start. >>>>> >>>>>I'm in the process of decoupling the scripts from our internal tooling - >>>>>I should be able to make available the Ambari stuff later this week. >>>>>Once >>>>>we merge the disparate forks of the Cisco codebase, there will be some >>>>>work to bring the playbooks up to date (i.e., use of Storm Flux), but >>>>>not >>>>>a ton. >>>>> >>>>>Mark >>>>> >>>>> >>>>> >>>>>On 12/9/15, 2:48 PM, "Bryan Taylor" <[email protected]> wrote: >>>>> >>>>>> >>>>>>Seems like a good list. I'm probably not your UI guy, but I'll be happy >>>>>>to >>>>>>help with any of the others. Test coverage is probably a good place for >>>>>>me >>>>>>to start learning my way around. Do we have CI set up? I see a travis >>>>>>file >>>>>>in the code. >>>>>> >>>>>>Regarding automation deployments, are we targeting Ambari or something >>>>>>else? >>>>>> >>>>>>On the hadoop component compatibility, I see from the opens-streaming >>>>>>pom >>>>>>that we are using Storm-0.9.2, Kafka 0.8.0, Hadoop 2.2.0, and HBase >>>>>>0.98.0-hadoop2. These are all several iterations old. How aggressive do >>>>>>we >>>>>>want to be, generally, with tracking new releases? Are Hive, Flume, and >>>>>>Spark also going to up rev for us? >>>>>> >>>>>>Bryan >>>>>> >>>>>>On 12/9/15 12:22 PM, "James Sirota" <[email protected]> wrote: >>>>>> >>>>>>>Hi Brian, >>>>>>> >>>>>>>Welcome. Glad to have you contribute. There will be changes to the >>>>>>>code >>>>>>>base that the community will contribute back shortly. We are waiting >>>>>>>for >>>>>>>the Jira to be setup so the backlog can be created and voted on. I >>>>>>>think >>>>>>>the overall feeling is that we need to make the code base compatible >>>>>>>with >>>>>>>the latest version of HDP, automate deployments, increase test >>>>>>>coverage, >>>>>>>and start working on a new UI. There may be more significant >>>>>>>architectural changes to the code base, but we need to get the >>>>>>>essential >>>>>>>items knocked out before we go there. >>>>>>> >>>>>>>Thanks, >>>>>>>James >>>>>>> >>>>>>> >>>>>>> >>>>>>> >>>>>>>On 12/9/15, 11:52 AM, "Bryan Taylor" <[email protected]> wrote: >>>>>>> >>>>>>>>Hi Folks, >>>>>>>> >>>>>>>>I just joined the list and thought I'd say "hi". I work at Rackspace >>>>>>>>and >>>>>>>>will be joining Andrew Hartnett's team there and hacking on metron. >>>>>>>>This >>>>>>>>is my first ASF project and I'm looking forward to being part of this >>>>>>>>community. >>>>>>>> >>>>>>>>I'm curious what the development vision is for metron. What do people >>>>>>>>like and not like about the codebase? I gather this code transitioned >>>>>>>>from a Cisco internal project and is incubating now at the ASF. Are >>>>>>>>there any code changes that need to be made to support incubation? >>>>>>>> >>>>>>>>Bryan >>>>>> >>>> >>>> >> >>
