John, Nice triage. After the nexus3 slowdown was fixed last week (routing change?) Casablanca comes up faster. Several of us like yourself are keenly actively bringing up Casablanca for vFW testing - it is good to compare notes. Thanks for the PV workaround - I find that deleting the pv/pvc and clearing the nfs directory for any pod in the middle of a -set *.enabled=false|true cycle fixes issues for those pods that normally cannot restart cleanly - as you mentioned. For pnda - I would keep this pod keyed off dcae disabled as it currently has a nodeport conflict that is being fixed - It has never been tested in a fully deployment - so the fact the image is wrong or missing makes sense. There is a separate thread on the db deadlock that needs a jira. Cross-threading here but for anyone seeing multiple config job restart failures - as long as one of them (the latest) eventually is 0/1 Completed - you are good for that pod - Mike is bringing in helm hooks as a better pattern. Many thanks for the 6 hour oof quicker restart verified option. Vvp is disabled by default - so no current CD deployment has tested it much - I do have a couple long-duration deployments but the eventual evictions that occur system wide - caused vvp to be lost in the noise - I thought it was resource starvation related - that triage is nice and merits a jira at least if the dns service name needs adjustment.
/michael From: [email protected] <[email protected]> On Behalf Of Lotoski, John Sent: Saturday, December 29, 2018 3:32 PM To: [email protected] Subject: [onap-discuss] Casablanca metal install lessons learned, plus 2 unresolved items (pnda + vvp) Hi All, I completed a Casablanca install on metal (3 G9 C7000 blades for K8s and 1 blade for Rancher + local Docker image caching proxy). In case others might find this information helpful, here are some notes on what I worked through followed by two items remaining that are unresolved: Due to ONAP Nexus low bandwidth and/or random or long docker image pull times, I needed to install a local docker image cache. This resolved a lot of the random issues and failures I was encountering. The local cache was setup per the following references: - Add Nexus proxy procedure for ONAP deployments - LF Nexus3 is timing out periodically: https://jira.onap.org/browse/LOG-355 - https://wiki.onap.org/display/DW/Cloud+Native+Deployment#CloudNativeDeployment-NexusProxy Once the local cache was in place, component-wise installation of ONAP modules helped isolate and troubleshoot remaining issues as per the following reference's approach: - https://wiki.onap.org/display/DW/Cloud+Native+Deployment#CloudNativeDeployment-CasablancaDeploymentExamples Example: sudo helm deploy onap local/onap --namespace onap -f onap/resources/environments/disable-allcharts.yaml --set global.repository=onapregistry:5000 --set dmaap.enabled=true The Dmaap module has an issue not yet specifically mentioned in a JIRA ticket (that I could find) which required a workaround applied to the dmaap db pod that is discussed in this mailing list at https://lists.onap.org/g/onap-discuss/topic/27483391: kubectl exec -ti onap-dmaap-dmaap-dr-db-<pod-identifier> -n onap -- bash mysql -u root SET PASSWORD FOR 'root'@'localhost' = PASSWORD('datarouter'); GRANT ALL ON *.* TO 'root'@'localhost' WITH GRANT OPTION; CREATE USER datarouter; SET PASSWORD FOR 'datarouter' = PASSWORD('datarouter'); CREATE database datarouter; GRANT ALL ON `datarouter`.* TO 'datarouter'; quit; exit Most ONAP modules were able to be installed or removed at will using the "--set <ONAP_module>.enabled=true/false" approach mentioned above. However, both OOF and dcaegen2 modules failed to re-deploy after removing with helm and showed the following helm error type: onap-dcaegen2 1 Thu Dec 27 18:18:43 2018 FAILED dcaegen2-3.0.0 onap It turns out that for at least these two modules (OOF and dcaegen2), the persistent volumes refuse to complete termination as shown by the following command: # kubectl get pv -n onap | grep -i dcaegen2 onap-dcaegen2-dcae-db-data0 1Gi RWO Retain Terminating onap/onap-dcaegen2-dcae-db-data-onap-dcaegen2-dcae-db-0 onap-dcaegen2-dcae-db-data 1d onap-dcaegen2-dcae-db-data1 1Gi RWO Retain Terminating onap/onap-dcaegen2-dcae-db-data-onap-dcaegen2-dcae-db-1 onap-dcaegen2-dcae-db-data 1d onap-dcaegen2-dcae-redis0 10Mi RWO Retain Terminating onap/onap-dcaegen2-dcae-redis-data-onap-dcaegen2-dcae-redis-1 manual 1d onap-dcaegen2-dcae-redis1 10Mi RWO Retain Terminating onap/onap-dcaegen2-dcae-redis-data-onap-dcaegen2-dcae-redis-3 manual 1d onap-dcaegen2-dcae-redis3 10Mi RWO Retain Terminating onap/onap-dcaegen2-dcae-redis-data-onap-dcaegen2-dcae-redis-4 manual 1d onap-dcaegen2-dcae-redis6 10Mi RWO Retain Terminating onap/onap-dcaegen2-dcae-redis-data-onap-dcaegen2-dcae-redis-5 manual 1d onap-dcaegen2-dcae-redis7 10Mi RWO Retain Terminating onap/onap-dcaegen2-dcae-redis-data-onap-dcaegen2-dcae-redis-2 manual 1d onap-dcaegen2-dcae-redis9 10Mi RWO Retain Terminating onap/onap-dcaegen2-dcae-redis-data-onap-dcaegen2-dcae-redis-0 manual 1d Termination can be allowed to complete by editing each of the pv resources which are stuck in the "Terminating" state and commenting out the "finalizer" lines: # kubectl edit persistentvolume <persistentvolumename> <...snip...> # Add the comment hash character to the two lines below: # finalizers: # - kubernetes.io/pv-protection <...snip...> There are also persistent volume claims which will show as "lost" as well with each associated module after the termination above completes. These can be deleted with: NAME STATUS VOLUME CAPACITY ACCESS MODES STORAGECLASS AGE onap-dcaegen2-dcae-db-data-onap-dcaegen2-dcae-db-0 Lost onap-dcaegen2-dcae-db-data0 0 onap-dcaegen2-dcae-db-data 1d <...snip...> kubectl -n onap delete pvc <persistentvolumeclaim> After clearing the stuck "Terminating" persistent volumes and optionally the "Lost" persistent volume claims and purging the corresponding /dockernfs-data/<onap_module> directory, these modules can be installed again using the "--set <ONAP_module>.enabled=true/false" approach. I did experience an occasional problem with the OOF pods which was traced to a database deadlock condition seen in the db-init logs from the onap-oof-oof-cmso-service pod: VM_ARGS= . ____ _ __ _ _ /\\ / ___'_ __ _ _(_)_ __ __ _ \ \ \ \ ( ( )\___ | '_ | '_| | '_ \/ _` | \ \ \ \ \\/ ___)| |_)| | | | | || (_| | ) ) ) ) ' |____| .__|_| |_|_| |_\__, | / / / / =========|_|==============|___/=/_/_/_/ :: Spring Boot :: (v2.0.6.RELEASE) 2018-12-25 23:06:59.310 INFO 8 --- [ main] o.o.o.c.liquibase.LiquibaseApplication : Starting LiquibaseApplication on onap-oof-oof-cmso-service-5f69655dcd-72fv6 with PID 8 (/opt/app/cmso-dbinit/app.jar started by root in /opt/app/cmso-dbinit) 2018-12-25 23:06:59.339 INFO 8 --- [ main] o.o.o.c.liquibase.LiquibaseApplication : No active profile set, falling back to default profiles: default 2018-12-25 23:06:59.415 INFO 8 --- [ main] s.c.a.AnnotationConfigApplicationContext : Refreshing org.springframework.context.annotation.AnnotationConfigApplicationContext@d44fc21<mailto:org.springframework.context.annotation.AnnotationConfigApplicationContext@d44fc21>: startup date [Tue Dec 25 23:06:59 UTC 2018]; root of context hierarchy 2018-12-25 23:07:00.633 INFO 8 --- [ main] com.zaxxer.hikari.HikariDataSource : HikariPool-1 - Starting... 2018-12-25 23:07:00.752 INFO 8 --- [ main] com.zaxxer.hikari.HikariDataSource : HikariPool-1 - Start completed. 2018-12-25 23:07:00.867 WARN 8 --- [ main] s.c.a.AnnotationConfigApplicationContext : Exception encountered during context initialization - cancelling refresh attempt: org.springframework.beans.factory.BeanCreationException: Error creating bean with name 'liquibase' defined in class path resource [org/onap/optf/cmso/liquibase/LiquibaseData.class]: Invocation of init method failed; nested exception is liquibase.exception.LockException: liquibase.exception.DatabaseException: liquibase.exception.DatabaseException: java.sql.SQLTransactionRollbackException: (conn=28) Deadlock found when trying to get lock; try restarting transaction 2018-12-25 23:07:00.868 INFO 8 --- [ main] com.zaxxer.hikari.HikariDataSource : HikariPool-1 - Shutdown initiated... 2018-12-25 23:07:02.041 INFO 8 --- [ main] com.zaxxer.hikari.HikariDataSource : HikariPool-1 - Shutdown completed. 2018-12-25 23:07:02.055 INFO 8 --- [ main] ConditionEvaluationReportLoggingListener : Error starting ApplicationContext. To display the conditions report re-run your application with 'debug' enabled. 2018-12-25 23:07:02.061 ERROR 8 --- [ main] o.s.boot.SpringApplication : Application run failed org.springframework.beans.factory.BeanCreationException: Error creating bean with name 'liquibase' defined in class path resource [org/onap/optf/cmso/liquibase/LiquibaseData.class]: Invocation of init method failed; nested exception is liquibase.exception.LockException: liquibase.exception.DatabaseException: liquibase.exception.DatabaseException: java.sql.SQLTransactionRollbackException: (conn=28) Deadlock found when trying to get lock; try restarting transaction at org.springframework.beans.factory.support.AbstractAutowireCapableBeanFactory.initializeBean(AbstractAutowireCapableBeanFactory.java:1694) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.beans.factory.support.AbstractAutowireCapableBeanFactory.doCreateBean(AbstractAutowireCapableBeanFactory.java:573) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.beans.factory.support.AbstractAutowireCapableBeanFactory.createBean(AbstractAutowireCapableBeanFactory.java:495) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.beans.factory.support.AbstractBeanFactory.lambda$doGetBean$0(AbstractBeanFactory.java:317) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.beans.factory.support.DefaultSingletonBeanRegistry.getSingleton(DefaultSingletonBeanRegistry.java:222) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.beans.factory.support.AbstractBeanFactory.doGetBean(AbstractBeanFactory.java:315) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.beans.factory.support.AbstractBeanFactory.getBean(AbstractBeanFactory.java:199) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.beans.factory.support.DefaultListableBeanFactory.preInstantiateSingletons(DefaultListableBeanFactory.java:759) ~[spring-beans-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.context.support.AbstractApplicationContext.finishBeanFactoryInitialization(AbstractApplicationContext.java:867) ~[spring-context-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.context.support.AbstractApplicationContext.refresh(AbstractApplicationContext.java:548) ~[spring-context-5.0.10.RELEASE.jar!/:5.0.10.RELEASE] at org.springframework.boot.SpringApplication.refresh(SpringApplication.java:754) [spring-boot-2.0.6.RELEASE.jar!/:2.0.6.RELEASE] at org.springframework.boot.SpringApplication.refreshContext(SpringApplication.java:386) [spring-boot-2.0.6.RELEASE.jar!/:2.0.6.RELEASE] at org.springframework.boot.SpringApplication.run(SpringApplication.java:307) [spring-boot-2.0.6.RELEASE.jar!/:2.0.6.RELEASE] at org.springframework.boot.SpringApplication.run(SpringApplication.java:1242) [spring-boot-2.0.6.RELEASE.jar!/:2.0.6.RELEASE] If this deadlock condition was left on its own to resolve, it was observed to resolve itself after about 5 to 6 hours. Otherwise, simply re-deploying the same helm command again re-spawned the onap-oof-oof-cmso-service pod and the issue resolved itself within several minutes. Deleting the pod directly and allowing Kubernetes to re-spawn it may have also worked although I didn't try that. It was observed that the ONAP sdc module created a few "Init:Error" pods during its deployment process. A log trace of these failed pods show a timeout of 10 minutes while waiting for the sdc-dcae-be pod to be ready. Since there is a "Completed" version of each of the failed pods, these appear to not be an issue and I believe they can be deleted. # kubectl -n onap logs onap-sdc-sdc-dcae-be-tools-w7427 -c sdc-dcae-be-init-readiness 2018-12-25 23:34:23,677 - INFO - Checking if sdc-dcae-be is ready 2018-12-25 23:34:24,132 - INFO - sdc-dcae-be is not ready. <... snip ...> 2018-12-25 23:44:25,537 - INFO - Checking if sdc-dcae-be is ready 2018-12-25 23:44:25,974 - INFO - sdc-dcae-be is not ready. 2018-12-25 23:44:25,976 - WARNING - timed out waiting for 'sdc-dcae-be' to be ready # watch -n 1 'kubectl get pods --all-namespaces -o=wide | grep -vi running' Every 5.0s: kubectl get pods --all-namespaces -o=wide | grep -vi running Sat Dec 29 19:59:50 2018 NAMESPACE NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE onap onap-sdc-sdc-be-config-backend-975b9 0/1 Init:Error 0 1d 10.42.226.2 c7k02bl04 <none> onap onap-sdc-sdc-be-config-backend-b2h7s 0/1 Completed 0 1d 10.42.152.132 c7k02bl03 <none> onap onap-sdc-sdc-cs-config-cassandra-bmr22 0/1 Completed 0 1d 10.42.10.121 c7k02bl02 <none> onap onap-sdc-sdc-dcae-be-tools-7bk87 0/1 Init:Error 0 1d 10.42.184.54 c7k02bl03 <none> onap onap-sdc-sdc-dcae-be-tools-czncn 0/1 Init:Error 0 1d 10.42.157.87 c7k02bl04 <none> onap onap-sdc-sdc-dcae-be-tools-gshsw 0/1 Completed 0 1d 10.42.81.19 c7k02bl03 <none> The ONAP contrib module often generated contrib-netbox-app-provisioning errors on the first deployment without a matching pod of the same name completing successfully. The error was from a shell script, provision.sh, returning an error code of '7': onap onap-contrib-netbox-app-provisioning-4kjgn 0/1 Error 0 5h 10.42.201.172 c7k02bl03 <none> onap onap-contrib-netbox-app-provisioning-6jxtq 0/1 Error 0 5h 10.42.183.209 c7k02bl02 <none> onap onap-contrib-netbox-app-provisioning-lc7wb 0/1 Error 0 5h 10.42.229.201 c7k02bl02 <none> onap onap-contrib-netbox-app-provisioning-v5jdz 0/1 Error 0 5h 10.42.132.244 c7k02bl04 <none> I'm not sure exactly what was failing here, but removing and re-adding this module one or more times would result in one of the onap-contrib-netbox-app-provisioning pods completing successfully. As of right now, I seem to have a functioning Casablanca install! Yay! There are two outstanding issues I see that have yet to be resolved. Any input on these items is appreciated: 1) The ONAP PNDA module doesn't install as the images it tries to pull don't appear to be in the Nexus repo yet. This image doesn't appear in the docker-manifest-casablanca.csv used by the docker_prepull.sh script either: Events: Type Reason Age From Message ---- ------ ---- ---- ------- Normal Scheduled 53m default-scheduler Successfully assigned onap/onap-pnda-dcae-pnda-mirror-58d689f947-pp5mh to c7k02bl04 Normal Pulling 53m (x3 over 53m) kubelet, c7k02bl04 pulling image "onapregistry:5000/onap/org.onap.dcaegen2.deployments.pnda-mirror-container:5.0.0" Warning Failed 53m (x3 over 53m) kubelet, c7k02bl04 Failed to pull image "onapregistry:5000/onap/org.onap.dcaegen2.deployments.pnda-mirror-container:5.0.0": rpc error: code = Unknown desc = Error: image onap/org.onap.dcaegen2.deployments.pnda-mirror-container:5.0.0 not found Warning Failed 53m (x3 over 53m) kubelet, c7k02bl04 Error: ErrImagePull Normal SandboxChanged 52m (x7 over 53m) kubelet, c7k02bl04 Pod sandbox changed, it will be killed and re-created. Warning Failed 8m (x203 over 53m) kubelet, c7k02bl04 Error: ImagePullBackOff Normal BackOff 3m (x224 over 53m) kubelet, c7k02bl04 Back-off pulling image "onapregistry:5000/onap/org.onap.dcaegen2.deployments.pnda-mirror-container:5.0.0" 2) The VVP imagescanner pod (onap-vvp-vvp-imagescanner-<pod_identifier>) is restarting a few times per hour. The following shows there is a timeout waiting for communication to a redis server that isn't happening: # kubectl -n onap logs onap-vvp-vvp-imagescanner-<pod_identifier> imagescanner-worker # # or # kubectl -n onap logs onap-vvp-vvp-imagescanner-<pod_identifier> notifications-worker # # results in: [2018-12-29 18:43:08,361: ERROR/MainProcess] consumer: Cannot connect to redis://redis:6379//: Error -2 connecting to redis:6379. Name does not resolve.. Trying again in 32.00 seconds... Thanks, John This email and the information contained herein is proprietary and confidential and subject to the Amdocs Email Terms of Service, which you may review at https://www.amdocs.com/about/email-terms-of-service <https://www.amdocs.com/about/email-terms-of-service> -=-=-=-=-=-=-=-=-=-=-=- Links: You receive all messages sent to this group. View/Reply Online (#14724): https://lists.onap.org/g/onap-discuss/message/14724 Mute This Topic: https://lists.onap.org/mt/28883772/21656 Group Owner: [email protected] Unsubscribe: https://lists.onap.org/g/onap-discuss/unsub [[email protected]] -=-=-=-=-=-=-=-=-=-=-=-
