Also if the problem is you are filling up /tmp inside your robot container try
commenting out this line inside your robot contaiener
/var/opt/ONAP/robot/resources/ test_templates/model_test_template.robot
#Download CSAR ${catalog_service_id}
This will skip the step of downloading the CSAR to the /tmp directory of your
robot container which if the csar is large could be the problem.
My robot continer is at 32% disk usage after doing 20 or so vFW creations but
yours is at 45% so you need to see why your robot container filesystem seems to
have a higher utilization.
Brian
From: FREEMAN, BRIAN D
Sent: Monday, March 30, 2020 8:54 AM
To: 'Rajendra P Mishra (RP)' <[email protected]>
Cc: '[email protected]' <[email protected]>; STARK, STEVEN
<[email protected]>
Subject: RE: [onap-discuss] [OVP][SO] Heatbridge failure while doing ovp
validation on reference VNF
My azure install was not a full onap install so you may want to turn off
optional components if you don’t have enough total ephemeral storage across all
your nodes.
Brian
From: FREEMAN, BRIAN D
Sent: Monday, March 30, 2020 8:53 AM
To: Rajendra P Mishra (RP)
<[email protected]<mailto:[email protected]>>
Cc: [email protected]<mailto:[email protected]>; STARK,
STEVEN <[email protected]<mailto:[email protected]>>
Subject: RE: [onap-discuss] [OVP][SO] Heatbridge failure while doing ovp
validation on reference VNF
Look at kubectl describe nodes and see how close to your ephemeral storage
capacity you are.
Some stats from my Azure environment.
root@onapdldk-robot-robot-577df6887c-njrc5:/# df -Th
Filesystem Type Size Used Avail
Use% Mounted on
overlay overlay 97G 31G 67G
32% /
tmpfs tmpfs 64M 0 64M
0% /dev
tmpfs tmpfs 14G 0 14G
0% /sys/fs/cgroup
tmpfs tmpfs 14G 158M 14G
2% /dev/shm
/dev/sda1 ext4 97G 31G 67G
32% /etc/localtime
173.163.0.4:/dockerdata-nfs/onapdldk-robot/robot/logs nfs4 497G 84G 413G
17% /share/logs
tmpfs tmpfs 14G 12K 14G
1% /run/secrets/kubernetes.io/serviceaccount
tmpfs tmpfs 14G 0 14G
0% /proc/acpi
tmpfs tmpfs 14G 0 14G
0% /proc/scsi
tmpfs tmpfs 14G 0 14G
0% /sys/firmware
bf1936@njcdtl01bf1936:~/ONAP/AZURE/ONAPDlDK/oom/kubernetes/robot$ kubectl -n
onap top nodes
NAME CPU(cores) CPU% MEMORY(bytes) MEMORY%
aks-nodepool1-97482552-0 883m 11% 18723Mi 78%
aks-nodepool1-97482552-1 4375m 55% 17451Mi 72%
aks-nodepool1-97482552-2 1849m 23% 12245Mi 51%
aks-nodepool1-97482552-3 1478m 18% 17252Mi 71%
aks-nodepool1-97482552-4 5382m 68% 21750Mi 90%
aks-nodepool1-97482552-5 4617m 59% 14628Mi 61%
aks-nodepool1-97482552-6 1179m 15% 19062Mi 79%
kubectl describe nodes
For node that robot is running:
Capacity:
attachable-volumes-azure-disk: 32
cpu: 8
ephemeral-storage: 101445900Ki
hugepages-1Gi: 0
hugepages-2Mi: 0
memory: 28789864Ki
pods: 110
Allocatable:
attachable-volumes-azure-disk: 32
cpu: 7820m
ephemeral-storage: 93492541286
hugepages-1Gi: 0
hugepages-2Mi: 0
memory: 24541288Ki
pods: 110
From: Rajendra P Mishra (RP)
<[email protected]<mailto:[email protected]>>
Sent: Sunday, March 29, 2020 12:06 PM
To: FREEMAN, BRIAN D <[email protected]<mailto:[email protected]>>
Cc: [email protected]<mailto:[email protected]>; STARK,
STEVEN <[email protected]<mailto:[email protected]>>
Subject: Re: [onap-discuss] [OVP][SO] Heatbridge failure while doing ovp
validation on reference VNF
NFS drive is 2TB
Here is the file system space usage in one of the node.
----
aarna@anod-master:~$ df -Th
Filesystem Type Size Used Avail Use% Mounted on
udev devtmpfs 16G 0 16G 0% /dev
tmpfs tmpfs 3.2G 1.3M 3.2G 1% /run
/dev/vda1 ext4 196G 87G 110G 45% /
tmpfs tmpfs 16G 0 16G 0% /dev/shm
tmpfs tmpfs 5.0M 0 5.0M 0% /run/lock
tmpfs tmpfs 16G 0 16G 0% /sys/fs/cgroup
/dev/vda15 vfat 105M 3.6M 101M 4% /boot/efi
10.240.71.179:/dockerdata-nfs nfs4 1.9T 23G 1.8T 2% /dockerdata-nfs
tmpfs tmpfs 3.2G 0 3.2G 0% /run/user/1000
---------
Disk usage inside robot is here.
-------
aarna@anod-master:~$ kubectl get pod -n onap | grep robot
dev-robot-robot-6d444f4cdd-rrwqj 0/1 Evicted
0 5h2m
dev-robot-robot-6d444f4cdd-sn5sl 1/1 Running
0 4h47m
aarna@anod-master:~$ kubectl exec -it -n onap dev-robot-robot-6d444f4cdd-sn5sl
-- /bin/bash
root@dev-robot-robot-6d444f4cdd-sn5sl:/# df -Th
Filesystem Type Size Used Avail
Use% Mounted on
overlay overlay 196G 88G 109G
45% /
tmpfs tmpfs 64M 0 64M
0% /dev
tmpfs tmpfs 16G 0 16G
0% /sys/fs/cgroup
/dev/vda1 ext4 196G 88G 109G
45% /etc/localtime
tmpfs tmpfs 16G 0 16G
0% /dev/shm
10.240.71.179:/dockerdata-nfs/dev-robot/robot/logs nfs4 1.9T 23G 1.8T
2% /share/logs
tmpfs tmpfs 16G 12K 16G
1%
/run/secrets/kubernetes.io/serviceaccount<https://urldefense.proofpoint.com/v2/url?u=http-3A__kubernetes.io_serviceaccount&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=Vk1goR9Pg3FUGQEz_PEKUp_E_Y4jBvPuiwnpGZijrNM&s=RLmzIwNVRpfgGrovLp8mIHGvUoruK9qxOUINwcTCuaA&e=>
tmpfs tmpfs 16G 0 16G
0% /proc/acpi
tmpfs tmpfs 16G 0 16G
0% /proc/scsi
tmpfs tmpfs 16G 0 16G
0% /sys/firmware
-----------
For my ONAP setup, each node is hosted on a VM. Each VM has around 196GB root
file system.
NFS is setup on a seperate server outside of these VMs.
I am trying to increase the size of the root file system from 196GB to 400GB.
What I do not understand is why 196GB is not enough ? and where is the space
getting consumed.
thanks,
-RP Mishra
On Sun, Mar 29, 2020 at 4:44 PM FREEMAN, BRIAN D
<[email protected]<mailto:[email protected]>> wrote:
I see now 27g out of 2tb it must be your k8 config is somehow different .
What does disk usage inside the robot container look like before it is evicted ?
When you created the k8 nodes how much storage did you give eack k8 ?
Maybe put robot on its own k8 host ?
Sent via the Samsung Galaxy S8, an AT&T 5G Evolution capable smartphone
-------- Original message --------
From: Rajendra Prasad Mishra
<[email protected]<mailto:[email protected]>>
Date: 3/29/20 6:39 AM (GMT-05:00)
To: "FREEMAN, BRIAN D" <[email protected]<mailto:[email protected]>>
Cc: [email protected]<mailto:[email protected]>, "STARK,
STEVEN" <[email protected]<mailto:[email protected]>>
Subject: Re: [onap-discuss] [OVP][SO] Heatbridge failure while doing ovp
validation on reference VNF
Hi Brian, Steve,
Deleting /share/logs and /dockerdata-nfs/dev-robot/robot/logs/ is not
helping. Robot pod is still getting killed due to space problems.
I had to re-install the whole ONAP to get rid of the problem. But it comes back
soon once I run some larger VNF (like vFW).
I tried increasing the storage allocated to robot in the values.yaml file
before installing ONAP, but I still get the same space problem when I run OVP
tests.
(below is the diff of the values.yaml file)
Is there a clean way to increase the "ephemeral-storage" for robot to some
large value so that I can do multiple runs of complex VNF without running into
space issues.
My /dockerdata-nfs is mounted on a 2TB SSD and only 21GB is used, so we have
enough space available.
---
aarna@anod-master:~/oom/kubernetes/robot$ git diff values.yaml
diff --git a/values.yaml b/values.yaml
index 3696057..9cda2c1 100644
--- a/values.yaml
+++ b/values.yaml
@@ -23,12 +23,12 @@ global: # global defaults
# application image
repository:
nexus3.onap.org:10001<https://urldefense.proofpoint.com/v2/url?u=http-3A__nexus3.onap.org-3A10001&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=ZlV4kOGHmpRjjv6cp1gx66F9Dv4gkllAc6ExmLAc6L4&e=>
image: onap/testsuite:1.5.4
-pullPolicy: Always
+pullPolicy: IfNotPresent
ubuntuInitImage: oomk8s/ubuntu-init:2.0.0
# flag to enable debugging - application support required
-debugEnabled: false
+debugEnabled: true
#################################################################
# Application configuration defaults.
@@ -254,6 +254,6 @@ persistence:
##
# storageClass: "-"
accessMode: ReadWriteMany
- size: 2Gi
+ size: 4Gi
mountPath: /dockerdata-nfs
mountSubPath: robot/logs
-----------------
Here is the pod that is killed.
-----------
aarna@anod-master:~$ kubectl describe pod -n onap
dev-robot-robot-6d444f4cdd-bkm4x
Name: dev-robot-robot-6d444f4cdd-bkm4x
Namespace: onap
Priority: 0
Node: dublin05/
Start Time: Fri, 27 Mar 2020 09:30:42 +0000
Labels: app=robot
pod-template-hash=6d444f4cdd
release=dev-robot
Annotations:
cni.projectcalico.org/podIP<https://urldefense.proofpoint.com/v2/url?u=http-3A__cni.projectcalico.org_podIP&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=4Ie1PTpTbBekcYyflcwb1Ppzy3DltykU2VI8WFIqGo4&e=>:
10.42.5.41/32<https://urldefense.proofpoint.com/v2/url?u=http-3A__10.42.5.41_32&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=PlvGS8KEgxKtPrUx0Fxgu1icZTwTwj_p31rCbEqFEVw&e=>
Status: Failed
Reason: Evicted
Message: The node was low on resource: ephemeral-storage. Container
robot was using 125078316Ki, which exceeds its request of 0.
IP:
Controlled By: ReplicaSet/dev-robot-robot-6d444f4cdd
Containers:
robot:
Image:
nexus3.onap.org:10001/onap/testsuite:1.5.4<https://urldefense.proofpoint.com/v2/url?u=http-3A__nexus3.onap.org-3A10001_onap_testsuite-3A1.5.4&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=uCMjhaGdBG8rdsixT8QAq7iFAB_mV4aWob5OD2Ou_is&e=>
Port: 88/TCP
Host Port: 0/TCP
Liveness: tcp-socket :88 delay=180s timeout=1s period=10s #success=1
#failure=3
Readiness: tcp-socket :88 delay=180s timeout=1s period=10s #success=1
#failure=3
Environment: <none>
Mounts:
/dev/shm from dshm (rw)
/etc/localtime from localtime (ro)
/share/config from robot-eteshare (rw)
/share/logs from robot-logs (rw)
/var/run/secrets/kubernetes.io/serviceaccount<https://urldefense.proofpoint.com/v2/url?u=http-3A__kubernetes.io_serviceaccount&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=62KOK4XwlhhxKcBn85bXtwlUakwbOIegXv08sO1GlqE&e=>
from default-token-qd76w (ro)
Volumes:
robot-logs:
Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim
in the same namespace)
ClaimName: dev-robot-robot
ReadOnly: false
dshm:
Type: EmptyDir (a temporary directory that shares a pod's lifetime)
Medium: Memory
SizeLimit: <unset>
localtime:
Type: HostPath (bare host directory volume)
Path: /etc/localtime
HostPathType:
robot-eteshare:
Type: ConfigMap (a volume populated by a ConfigMap)
Name: dev-robot-robot-eteshare-configmap
Optional: false
default-token-qd76w:
Type: Secret (a volume populated by a Secret)
SecretName: default-token-qd76w
Optional: false
QoS Class: BestEffort
Node-Selectors: <none>
Tolerations:
node.kubernetes.io/not-ready:NoExecute<https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_not-2Dready-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=5IOJYuAKepGEJxmwD8wQHsOePlpnB2x3gbU9Y5XCoIE&e=>
for 300s
node.kubernetes.io/unreachable:NoExecute<https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_unreachable-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=ujPOfc7-AgMLBQA_u0ZbKoJ-ZOlVJU4bhrOnx_O9wQc&s=cNdW_67uaUUVC_JSvX41xLh18gDEnPp4p7za_daCxH0&e=>
for 300s
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning Evicted 25m kubelet, dublin05 The node was low on resource:
ephemeral-storage. Container robot was using 125078316Ki, which exceeds its
request of 0.
Normal Killing 25m kubelet, dublin05 Killing container with id
docker://robot:Need to kill Pod
------------------------
thanks for your help,
-RP Mishra
On Mon, Mar 23, 2020 at 7:07 PM FREEMAN, BRIAN D
<[email protected]<mailto:[email protected]>> wrote:
IF you have been doing a lot of testing the robot file system can get pretty
large (500Meg per test in some cases)
Cleanup /share/logs inside the container (that is log files stored on the nfs)
and /tmp/csar where the distributed models are stored for use in testing like
vCPE
Brian
From: [email protected]<mailto:[email protected]>
<[email protected]<mailto:[email protected]>> On Behalf Of
Rajendra Prasad Mishra
Sent: Monday, March 23, 2020 9:25 AM
To: STARK, STEVEN <[email protected]<mailto:[email protected]>>
Cc: [email protected]<mailto:[email protected]>
Subject: Re: [onap-discuss] [OVP][SO] Heatbridge failure while doing ovp
validation on reference VNF
Thanks Steve, that worked. I no longer see the openstack authentication issue.
However, I am seeing the robot storage issues. When I run the validation test,
robot pod is killed because it is consuming more "ephemeral-storage" then
allocated.
Do you know how to increase the " ephemeral-storage" for robot pod ? This is a
standard El Alto installation with no change in default configuration.
Why does validation tests take up so much of space (127421328Ki)
-----------------
aarna@anod-master:~/oom/kubernetes/robot$ kubectl get pod -n onap | grep robot
dev-robot-robot-6d444f4cdd-94kkk 1/1 Running
0 3m49s
dev-robot-robot-6d444f4cdd-vspc9 0/1 Evicted
0 3d15h
aarna@anod-master:~/oom/kubernetes/robot$ kubectl describe pod -n onap
dev-robot-robot-6d444f4cdd-vspc9
Name: dev-robot-robot-6d444f4cdd-vspc9
Namespace: onap
Priority: 0
Node: dublin06/
Start Time: Thu, 19 Mar 2020 15:24:57 +0000
Labels: app=robot
pod-template-hash=6d444f4cdd
release=dev-robot
Annotations:
cni.projectcalico.org/podIP<https://urldefense.proofpoint.com/v2/url?u=http-3A__cni.projectcalico.org_podIP&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=MLwHGSQ5pPlDSCtDzE_p7lM5vcEUhSDD8zDoC3S4vn4&e=>:
10.42.1.56/32<https://urldefense.proofpoint.com/v2/url?u=http-3A__10.42.1.56_32&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=Pbnhw9Cxfa2A6klTNVsFKrsskeFRF0guKda-qrkjVC4&e=>
Status: Failed
Reason: Evicted
Message: The node was low on resource: ephemeral-storage. Container
robot was using 127421328Ki, which exceeds its request of 0.
IP:
Controlled By: ReplicaSet/dev-robot-robot-6d444f4cdd
Containers:
robot:
Image:
nexus3.onap.org:10001/onap/testsuite:1.5.4<https://urldefense.proofpoint.com/v2/url?u=http-3A__nexus3.onap.org-3A10001_onap_testsuite-3A1.5.4&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=zGRkR7bU0ERGCcc4ahXTnuiECXcRN__EkBCF_EOPrsM&e=>
Port: 88/TCP
Host Port: 0/TCP
Liveness: tcp-socket :88 delay=180s timeout=1s period=10s #success=1
#failure=3
Readiness: tcp-socket :88 delay=180s timeout=1s period=10s #success=1
#failure=3
Environment: <none>
Mounts:
/dev/shm from dshm (rw)
/etc/localtime from localtime (ro)
/share/config from robot-eteshare (rw)
/share/logs from robot-logs (rw)
/var/run/secrets/kubernetes.io/serviceaccount<https://urldefense.proofpoint.com/v2/url?u=http-3A__kubernetes.io_serviceaccount&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=FAoROrnwGV5yFeLt7ZfIM_LtdbNGb7eMT0YJCiD7XLw&e=>
from default-token-2mz7g (ro)
Volumes:
robot-logs:
Type: PersistentVolumeClaim (a reference to a PersistentVolumeClaim
in the same namespace)
ClaimName: dev-robot-robot
ReadOnly: false
dshm:
Type: EmptyDir (a temporary directory that shares a pod's lifetime)
Medium: Memory
SizeLimit: <unset>
localtime:
Type: HostPath (bare host directory volume)
Path: /etc/localtime
HostPathType:
robot-eteshare:
Type: ConfigMap (a volume populated by a ConfigMap)
Name: dev-robot-robot-eteshare-configmap
Optional: false
default-token-2mz7g:
Type: Secret (a volume populated by a Secret)
SecretName: default-token-2mz7g
Optional: false
QoS Class: BestEffort
Node-Selectors: <none>
Tolerations:
node.kubernetes.io/not-ready:NoExecute<https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_not-2Dready-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=0b-LBWz7xZSxyHD54a1S970SRnyV92DCFb8OC5hGa6g&e=>
for 300s
node.kubernetes.io/unreachable:NoExecute<https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_unreachable-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=lUcFt1b-6535RBbsO7aF5SFumNsJvLXBjdo1eIXnxgA&e=>
for 300s
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning Evicted 4m26s kubelet, dublin06 The node was low on resource:
ephemeral-storage. Container robot was using 127421328Ki, which exceeds its
request of 0.
Normal Killing 4m24s kubelet, dublin06 Killing container with id
docker://robot:Need to kill Pod
----------------------------------
thanks,
-RPMishra
On Fri, Mar 20, 2020 at 9:42 PM Steve Stark
<[email protected]<mailto:[email protected]>> wrote:
Great that helps. It looks like instantiation worked, the part that's failing
is some custom heat stack validation after the instantiation.
I was able to replicate the failure you're getting. Can you change the
admin_tenant value for your identity service to the tenant name instead of the
tenant ID and try again? I think in your case the tenant name = admin.
This is in the mariadb pod -
MariaDB [catalogdb]> update identity_services set admin_tenant="<tenant name>"
where id="<identity service id>";
Thanks
Steven
-=-=-=-=-=-=-=-=-=-=-=-
Links: You receive all messages sent to this group.
View/Reply Online (#20421): https://lists.onap.org/g/onap-discuss/message/20421
Mute This Topic: https://lists.onap.org/mt/72026138/21656
Group Owner: [email protected]
Unsubscribe: https://lists.onap.org/g/onap-discuss/unsub
[[email protected]]
-=-=-=-=-=-=-=-=-=-=-=-