NFS drive is 2TB Here is the file system space usage in one of the node. ---- aarna@anod-master:~$ df -Th Filesystem Type Size Used Avail Use% Mounted on udev devtmpfs 16G 0 16G 0% /dev tmpfs tmpfs 3.2G 1.3M 3.2G 1% /run /dev/vda1 ext4 196G 87G 110G 45% / tmpfs tmpfs 16G 0 16G 0% /dev/shm tmpfs tmpfs 5.0M 0 5.0M 0% /run/lock tmpfs tmpfs 16G 0 16G 0% /sys/fs/cgroup /dev/vda15 vfat 105M 3.6M 101M 4% /boot/efi 10.240.71.179:/dockerdata-nfs nfs4 1.9T 23G 1.8T 2% /dockerdata-nfs tmpfs tmpfs 3.2G 0 3.2G 0% /run/user/1000 ---------
Disk usage inside robot is here. ------- aarna@anod-master:~$ kubectl get pod -n onap | grep robot dev-robot-robot-6d444f4cdd-rrwqj 0/1 Evicted 0 5h2m dev-robot-robot-6d444f4cdd-sn5sl 1/1 Running 0 4h47m aarna@anod-master:~$ kubectl exec -it -n onap dev-robot-robot-6d444f4cdd-sn5sl -- /bin/bash root@dev-robot-robot-6d444f4cdd-sn5sl:/# df -Th Filesystem Type Size Used Avail Use% Mounted on overlay overlay 196G 88G 109G 45% / tmpfs tmpfs 64M 0 64M 0% /dev tmpfs tmpfs 16G 0 16G 0% /sys/fs/cgroup /dev/vda1 ext4 196G 88G 109G 45% /etc/localtime tmpfs tmpfs 16G 0 16G 0% /dev/shm 10.240.71.179:/dockerdata-nfs/dev-robot/robot/logs nfs4 1.9T 23G 1.8T 2% /share/logs tmpfs tmpfs 16G 12K 16G 1% /run/secrets/kubernetes.io/serviceaccount tmpfs tmpfs 16G 0 16G 0% /proc/acpi tmpfs tmpfs 16G 0 16G 0% /proc/scsi tmpfs tmpfs 16G 0 16G 0% /sys/firmware ----------- For my ONAP setup, each node is hosted on a VM. Each VM has around 196GB root file system. NFS is setup on a seperate server outside of these VMs. I am trying to increase the size of the root file system from 196GB to 400GB. What I do not understand is why 196GB is not enough ? and where is the space getting consumed. thanks, -RP Mishra On Sun, Mar 29, 2020 at 4:44 PM FREEMAN, BRIAN D <[email protected]> wrote: > I see now 27g out of 2tb it must be your k8 config is somehow different . > > What does disk usage inside the robot container look like before it is > evicted ? > > When you created the k8 nodes how much storage did you give eack k8 ? > > Maybe put robot on its own k8 host ? > > > Sent via the Samsung Galaxy S8, an AT&T 5G Evolution capable smartphone > > > > -------- Original message -------- > From: Rajendra Prasad Mishra <[email protected]> > Date: 3/29/20 6:39 AM (GMT-05:00) > To: "FREEMAN, BRIAN D" <[email protected]> > Cc: [email protected], "STARK, STEVEN" <[email protected]> > Subject: Re: [onap-discuss] [OVP][SO] Heatbridge failure while doing ovp > validation on reference VNF > > Hi Brian, Steve, > Deleting /share/logs and /dockerdata-nfs/dev-robot/robot/logs/ is > not helping. Robot pod is still getting killed due to space problems. > > I had to re-install the whole ONAP to get rid of the problem. But it comes > back soon once I run some larger VNF (like vFW). > > I tried increasing the storage allocated to robot in the values.yaml file > before installing ONAP, but I still get the same space problem when I run > OVP tests. > (below is the diff of the values.yaml file) > > Is there a clean way to increase the "*ephemeral-storage*" for *robot *to > some large value so that I can do multiple runs of complex VNF > without running into space issues. > My /dockerdata-nfs is mounted on a 2TB SSD and only 21GB is used, so we > have enough space available. > > --- > aarna@anod-master:~/oom/kubernetes/robot$ git diff values.yaml > diff --git a/values.yaml b/values.yaml > index 3696057..9cda2c1 100644 > --- a/values.yaml > +++ b/values.yaml > @@ -23,12 +23,12 @@ global: # global defaults > # application image > repository: nexus3.onap.org:10001 > <https://urldefense.proofpoint.com/v2/url?u=http-3A__nexus3.onap.org-3A10001&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=ZlV4kOGHmpRjjv6cp1gx66F9Dv4gkllAc6ExmLAc6L4&e=> > image: onap/testsuite:1.5.4 > -pullPolicy: Always > +pullPolicy: IfNotPresent > > ubuntuInitImage: oomk8s/ubuntu-init:2.0.0 > > # flag to enable debugging - application support required > -debugEnabled: false > +debugEnabled: true > > ################################################################# > # Application configuration defaults. > @@ -254,6 +254,6 @@ persistence: > ## > # storageClass: "-" > accessMode: ReadWriteMany > - size: 2Gi > + size: 4Gi > mountPath: /dockerdata-nfs > mountSubPath: robot/logs > ----------------- > > Here is the pod that is killed. > > ----------- > aarna@anod-master:~$ kubectl describe pod -n onap > dev-robot-robot-6d444f4cdd-bkm4x > Name: dev-robot-robot-6d444f4cdd-bkm4x > Namespace: onap > Priority: 0 > Node: dublin05/ > Start Time: Fri, 27 Mar 2020 09:30:42 +0000 > Labels: app=robot > pod-template-hash=6d444f4cdd > release=dev-robot > Annotations: cni.projectcalico.org/podIP > <https://urldefense.proofpoint.com/v2/url?u=http-3A__cni.projectcalico.org_podIP&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=4Ie1PTpTbBekcYyflcwb1Ppzy3DltykU2VI8WFIqGo4&e=>: > 10.42.5.41/32 > <https://urldefense.proofpoint.com/v2/url?u=http-3A__10.42.5.41_32&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=PlvGS8KEgxKtPrUx0Fxgu1icZTwTwj_p31rCbEqFEVw&e=> > Status: Failed > Reason: Evicted > Message: The node was low on resource: ephemeral-storage. Container > robot was using 125078316Ki, which exceeds its request of 0. > IP: > Controlled By: ReplicaSet/dev-robot-robot-6d444f4cdd > Containers: > robot: > Image: nexus3.onap.org:10001/onap/testsuite:1.5.4 > <https://urldefense.proofpoint.com/v2/url?u=http-3A__nexus3.onap.org-3A10001_onap_testsuite-3A1.5.4&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=uCMjhaGdBG8rdsixT8QAq7iFAB_mV4aWob5OD2Ou_is&e=> > Port: 88/TCP > Host Port: 0/TCP > Liveness: tcp-socket :88 delay=180s timeout=1s period=10s > #success=1 #failure=3 > Readiness: tcp-socket :88 delay=180s timeout=1s period=10s > #success=1 #failure=3 > Environment: <none> > Mounts: > /dev/shm from dshm (rw) > /etc/localtime from localtime (ro) > /share/config from robot-eteshare (rw) > /share/logs from robot-logs (rw) > /var/run/secrets/kubernetes.io/serviceaccount > <https://urldefense.proofpoint.com/v2/url?u=http-3A__kubernetes.io_serviceaccount&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=62KOK4XwlhhxKcBn85bXtwlUakwbOIegXv08sO1GlqE&e=> > from default-token-qd76w (ro) > Volumes: > robot-logs: > Type: PersistentVolumeClaim (a reference to a > PersistentVolumeClaim in the same namespace) > ClaimName: dev-robot-robot > ReadOnly: false > dshm: > Type: EmptyDir (a temporary directory that shares a pod's > lifetime) > Medium: Memory > SizeLimit: <unset> > localtime: > Type: HostPath (bare host directory volume) > Path: /etc/localtime > HostPathType: > robot-eteshare: > Type: ConfigMap (a volume populated by a ConfigMap) > Name: dev-robot-robot-eteshare-configmap > Optional: false > default-token-qd76w: > Type: Secret (a volume populated by a Secret) > SecretName: default-token-qd76w > Optional: false > QoS Class: BestEffort > Node-Selectors: <none> > Tolerations: node.kubernetes.io/not-ready:NoExecute > <https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_not-2Dready-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=5IOJYuAKepGEJxmwD8wQHsOePlpnB2x3gbU9Y5XCoIE&e=> > for 300s > node.kubernetes.io/unreachable:NoExecute > <https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_unreachable-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=cNdW_67uaUUVC_JSvX41xLh18gDEnPp4p7za_daCxH0&e=> > for 300s > Events: > Type Reason Age From Message > ---- ------ ---- ---- ------- > Warning Evicted 25m kubelet, dublin05 The node was low on resource: > ephemeral-storage. Container robot was using 125078316Ki, which exceeds its > request of 0. > Normal Killing 25m kubelet, dublin05 Killing container with id > docker://robot:Need to kill Pod > ------------------------ > > thanks for your help, > -RP Mishra > > On Mon, Mar 23, 2020 at 7:07 PM FREEMAN, BRIAN D <[email protected]> wrote: > >> IF you have been doing a lot of testing the robot file system can get >> pretty large (500Meg per test in some cases) >> >> >> >> Cleanup /share/logs inside the container (that is log files stored on the >> nfs) and /tmp/csar where the distributed models are stored for use in >> testing like vCPE >> >> >> >> Brian >> >> >> >> >> >> *From:* [email protected] <[email protected]> *On >> Behalf Of *Rajendra Prasad Mishra >> *Sent:* Monday, March 23, 2020 9:25 AM >> *To:* STARK, STEVEN <[email protected]> >> *Cc:* [email protected] >> *Subject:* Re: [onap-discuss] [OVP][SO] Heatbridge failure while doing >> ovp validation on reference VNF >> >> >> >> Thanks Steve, that worked. I no longer see the openstack authentication >> issue. >> >> >> >> However, I am seeing the robot storage issues. When I run the validation >> test, robot pod is killed because it is consuming more "ephemeral-storage" >> then allocated. >> >> Do you know how to increase the " ephemeral-storage" for robot pod ? This >> is a standard El Alto installation with no change in default configuration. >> >> Why does validation tests take up so much of space (127421328Ki) >> >> >> >> ----------------- >> >> aarna@anod-master:~/oom/kubernetes/robot$ kubectl get pod -n onap | grep >> robot >> dev-robot-robot-6d444f4cdd-94kkk 1/1 >> Running 0 3m49s >> dev-robot-robot-6d444f4cdd-vspc9 0/1 >> Evicted 0 3d15h >> aarna@anod-master:~/oom/kubernetes/robot$ kubectl describe pod -n onap >> dev-robot-robot-6d444f4cdd-vspc9 >> Name: dev-robot-robot-6d444f4cdd-vspc9 >> Namespace: onap >> Priority: 0 >> Node: dublin06/ >> Start Time: Thu, 19 Mar 2020 15:24:57 +0000 >> Labels: app=robot >> pod-template-hash=6d444f4cdd >> release=dev-robot >> Annotations: cni.projectcalico.org/podIP >> <https://urldefense.proofpoint.com/v2/url?u=http-3A__cni.projectcalico.org_podIP&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=MLwHGSQ5pPlDSCtDzE_p7lM5vcEUhSDD8zDoC3S4vn4&e=>: >> 10.42.1.56/32 >> <https://urldefense.proofpoint.com/v2/url?u=http-3A__10.42.1.56_32&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=Pbnhw9Cxfa2A6klTNVsFKrsskeFRF0guKda-qrkjVC4&e=> >> Status: Failed >> Reason: Evicted >> Message: * The node was low on resource: ephemeral-storage. >> Container robot was using 127421328Ki, which exceeds its request of 0.* >> IP: >> Controlled By: ReplicaSet/dev-robot-robot-6d444f4cdd >> Containers: >> robot: >> Image: nexus3.onap.org:10001/onap/testsuite:1.5.4 >> <https://urldefense.proofpoint.com/v2/url?u=http-3A__nexus3.onap.org-3A10001_onap_testsuite-3A1.5.4&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=zGRkR7bU0ERGCcc4ahXTnuiECXcRN__EkBCF_EOPrsM&e=> >> Port: 88/TCP >> Host Port: 0/TCP >> Liveness: tcp-socket :88 delay=180s timeout=1s period=10s >> #success=1 #failure=3 >> Readiness: tcp-socket :88 delay=180s timeout=1s period=10s >> #success=1 #failure=3 >> Environment: <none> >> Mounts: >> /dev/shm from dshm (rw) >> /etc/localtime from localtime (ro) >> /share/config from robot-eteshare (rw) >> /share/logs from robot-logs (rw) >> /var/run/secrets/kubernetes.io/serviceaccount >> <https://urldefense.proofpoint.com/v2/url?u=http-3A__kubernetes.io_serviceaccount&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=FAoROrnwGV5yFeLt7ZfIM_LtdbNGb7eMT0YJCiD7XLw&e=> >> from default-token-2mz7g (ro) >> Volumes: >> robot-logs: >> Type: PersistentVolumeClaim (a reference to a >> PersistentVolumeClaim in the same namespace) >> ClaimName: dev-robot-robot >> ReadOnly: false >> dshm: >> Type: EmptyDir (a temporary directory that shares a pod's >> lifetime) >> Medium: Memory >> SizeLimit: <unset> >> localtime: >> Type: HostPath (bare host directory volume) >> Path: /etc/localtime >> HostPathType: >> robot-eteshare: >> Type: ConfigMap (a volume populated by a ConfigMap) >> Name: dev-robot-robot-eteshare-configmap >> Optional: false >> default-token-2mz7g: >> Type: Secret (a volume populated by a Secret) >> SecretName: default-token-2mz7g >> Optional: false >> QoS Class: BestEffort >> Node-Selectors: <none> >> Tolerations: node.kubernetes.io/not-ready:NoExecute >> <https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_not-2Dready-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=0b-LBWz7xZSxyHD54a1S970SRnyV92DCFb8OC5hGa6g&e=> >> for 300s >> node.kubernetes.io/unreachable:NoExecute >> <https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_unreachable-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=lUcFt1b-6535RBbsO7aF5SFumNsJvLXBjdo1eIXnxgA&e=> >> for 300s >> Events: >> Type Reason Age From Message >> ---- ------ ---- ---- ------- >> >> * Warning Evicted 4m26s kubelet, dublin06 The node was low on >> resource: ephemeral-storage. Container robot was using 127421328Ki, which >> exceeds its request of 0. Normal Killing 4m24s kubelet, dublin06 >> Killing container with id docker://robot:Need to kill Pod* >> >> ---------------------------------- >> >> >> >> thanks, >> >> -RPMishra >> >> >> >> >> >> On Fri, Mar 20, 2020 at 9:42 PM Steve Stark <[email protected]> wrote: >> >> Great that helps. It looks like instantiation worked, the part that's >> failing is some custom heat stack validation after the instantiation. >> >> I was able to replicate the failure you're getting. Can you change the >> admin_tenant value for your identity service to the tenant name instead of >> the tenant ID and try again? I think in your case the tenant name = admin. >> >> This is in the mariadb pod - >> >> MariaDB [catalogdb]> update identity_services set admin_tenant="<tenant >> name>" where id="<identity service id>"; >> >> Thanks >> Steven >> >> > -=-=-=-=-=-=-=-=-=-=-=- Links: You receive all messages sent to this group. View/Reply Online (#20417): https://lists.onap.org/g/onap-discuss/message/20417 Mute This Topic: https://lists.onap.org/mt/72026138/21656 Group Owner: [email protected] Unsubscribe: https://lists.onap.org/g/onap-discuss/unsub [[email protected]] -=-=-=-=-=-=-=-=-=-=-=-
