Hi Brian, Steve,
     Deleting /share/logs and  /dockerdata-nfs/dev-robot/robot/logs/ is not
helping. Robot pod is still getting killed due to space problems.

I had to re-install the whole ONAP to get rid of the problem. But it comes
back soon once I run some larger VNF (like vFW).

I tried increasing the storage allocated to robot in the values.yaml file
before installing ONAP, but I still get the same space problem when I run
OVP tests.
(below is the diff of the values.yaml file)

Is there a clean way to increase the "*ephemeral-storage*" for *robot *to
some large value so that I can do multiple runs of complex VNF
without running into space issues.
My /dockerdata-nfs is mounted on a 2TB SSD and only 21GB is used, so we
have enough space available.

---
aarna@anod-master:~/oom/kubernetes/robot$ git diff  values.yaml
diff --git a/values.yaml b/values.yaml
index 3696057..9cda2c1 100644
--- a/values.yaml
+++ b/values.yaml
@@ -23,12 +23,12 @@ global: # global defaults
 # application image
 repository: nexus3.onap.org:10001
 image: onap/testsuite:1.5.4
-pullPolicy: Always
+pullPolicy: IfNotPresent

 ubuntuInitImage: oomk8s/ubuntu-init:2.0.0

 # flag to enable debugging - application support required
-debugEnabled: false
+debugEnabled: true

 #################################################################
 # Application configuration defaults.
@@ -254,6 +254,6 @@ persistence:
   ##
   # storageClass: "-"
   accessMode: ReadWriteMany
-  size: 2Gi
+  size: 4Gi
   mountPath: /dockerdata-nfs
   mountSubPath: robot/logs
-----------------

Here is the pod that is killed.

-----------
aarna@anod-master:~$ kubectl describe pod -n onap
dev-robot-robot-6d444f4cdd-bkm4x
Name:           dev-robot-robot-6d444f4cdd-bkm4x
Namespace:      onap
Priority:       0
Node:           dublin05/
Start Time:     Fri, 27 Mar 2020 09:30:42 +0000
Labels:         app=robot
                pod-template-hash=6d444f4cdd
                release=dev-robot
Annotations:    cni.projectcalico.org/podIP: 10.42.5.41/32
Status:         Failed
Reason:         Evicted
Message:        The node was low on resource: ephemeral-storage. Container
robot was using 125078316Ki, which exceeds its request of 0.
IP:
Controlled By:  ReplicaSet/dev-robot-robot-6d444f4cdd
Containers:
  robot:
    Image:        nexus3.onap.org:10001/onap/testsuite:1.5.4
    Port:         88/TCP
    Host Port:    0/TCP
    Liveness:     tcp-socket :88 delay=180s timeout=1s period=10s
#success=1 #failure=3
    Readiness:    tcp-socket :88 delay=180s timeout=1s period=10s
#success=1 #failure=3
    Environment:  <none>
    Mounts:
      /dev/shm from dshm (rw)
      /etc/localtime from localtime (ro)
      /share/config from robot-eteshare (rw)
      /share/logs from robot-logs (rw)
      /var/run/secrets/kubernetes.io/serviceaccount from
default-token-qd76w (ro)
Volumes:
  robot-logs:
    Type:       PersistentVolumeClaim (a reference to a
PersistentVolumeClaim in the same namespace)
    ClaimName:  dev-robot-robot
    ReadOnly:   false
  dshm:
    Type:       EmptyDir (a temporary directory that shares a pod's
lifetime)
    Medium:     Memory
    SizeLimit:  <unset>
  localtime:
    Type:          HostPath (bare host directory volume)
    Path:          /etc/localtime
    HostPathType:
  robot-eteshare:
    Type:      ConfigMap (a volume populated by a ConfigMap)
    Name:      dev-robot-robot-eteshare-configmap
    Optional:  false
  default-token-qd76w:
    Type:        Secret (a volume populated by a Secret)
    SecretName:  default-token-qd76w
    Optional:    false
QoS Class:       BestEffort
Node-Selectors:  <none>
Tolerations:     node.kubernetes.io/not-ready:NoExecute for 300s
                 node.kubernetes.io/unreachable:NoExecute for 300s
Events:
  Type     Reason   Age   From               Message
  ----     ------   ----  ----               -------
  Warning  Evicted  25m   kubelet, dublin05  The node was low on resource:
ephemeral-storage. Container robot was using 125078316Ki, which exceeds its
request of 0.
  Normal   Killing  25m   kubelet, dublin05  Killing container with id
docker://robot:Need to kill Pod
------------------------

thanks for your help,
-RP Mishra

On Mon, Mar 23, 2020 at 7:07 PM FREEMAN, BRIAN D <[email protected]> wrote:

> IF you have been doing a lot of testing the robot file system can get
> pretty large (500Meg per test in some cases)
>
>
>
> Cleanup /share/logs inside the container (that is log files stored on the
> nfs) and /tmp/csar where the distributed models are stored for use in
> testing like vCPE
>
>
>
> Brian
>
>
>
>
>
> *From:* [email protected] <[email protected]> *On
> Behalf Of *Rajendra Prasad Mishra
> *Sent:* Monday, March 23, 2020 9:25 AM
> *To:* STARK, STEVEN <[email protected]>
> *Cc:* [email protected]
> *Subject:* Re: [onap-discuss] [OVP][SO] Heatbridge failure while doing
> ovp validation on reference VNF
>
>
>
> Thanks Steve, that worked. I no longer see the openstack authentication
> issue.
>
>
>
> However, I am seeing the robot storage issues. When I run the validation
> test, robot pod is killed because it is consuming more  "ephemeral-storage"
> then allocated.
>
> Do you know how to increase the " ephemeral-storage" for robot pod ? This
> is a standard El Alto installation with no change in default configuration.
>
> Why does validation tests take up so much of space (127421328Ki)
>
>
>
> -----------------
>
> aarna@anod-master:~/oom/kubernetes/robot$ kubectl get pod -n onap | grep
> robot
> dev-robot-robot-6d444f4cdd-94kkk                              1/1
> Running      0          3m49s
> dev-robot-robot-6d444f4cdd-vspc9                              0/1
> Evicted      0          3d15h
> aarna@anod-master:~/oom/kubernetes/robot$ kubectl describe pod -n onap
> dev-robot-robot-6d444f4cdd-vspc9
> Name:           dev-robot-robot-6d444f4cdd-vspc9
> Namespace:      onap
> Priority:       0
> Node:           dublin06/
> Start Time:     Thu, 19 Mar 2020 15:24:57 +0000
> Labels:         app=robot
>                 pod-template-hash=6d444f4cdd
>                 release=dev-robot
> Annotations:    cni.projectcalico.org/podIP
> <https://urldefense.proofpoint.com/v2/url?u=http-3A__cni.projectcalico.org_podIP&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=MLwHGSQ5pPlDSCtDzE_p7lM5vcEUhSDD8zDoC3S4vn4&e=>:
> 10.42.1.56/32
> <https://urldefense.proofpoint.com/v2/url?u=http-3A__10.42.1.56_32&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=Pbnhw9Cxfa2A6klTNVsFKrsskeFRF0guKda-qrkjVC4&e=>
> Status:         Failed
> Reason:         Evicted
> Message:      *  The node was low on resource: ephemeral-storage.
> Container robot was using 127421328Ki, which exceeds its request of 0.*
> IP:
> Controlled By:  ReplicaSet/dev-robot-robot-6d444f4cdd
> Containers:
>   robot:
>     Image:        nexus3.onap.org:10001/onap/testsuite:1.5.4
> <https://urldefense.proofpoint.com/v2/url?u=http-3A__nexus3.onap.org-3A10001_onap_testsuite-3A1.5.4&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=zGRkR7bU0ERGCcc4ahXTnuiECXcRN__EkBCF_EOPrsM&e=>
>     Port:         88/TCP
>     Host Port:    0/TCP
>     Liveness:     tcp-socket :88 delay=180s timeout=1s period=10s
> #success=1 #failure=3
>     Readiness:    tcp-socket :88 delay=180s timeout=1s period=10s
> #success=1 #failure=3
>     Environment:  <none>
>     Mounts:
>       /dev/shm from dshm (rw)
>       /etc/localtime from localtime (ro)
>       /share/config from robot-eteshare (rw)
>       /share/logs from robot-logs (rw)
>       /var/run/secrets/kubernetes.io/serviceaccount
> <https://urldefense.proofpoint.com/v2/url?u=http-3A__kubernetes.io_serviceaccount&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=FAoROrnwGV5yFeLt7ZfIM_LtdbNGb7eMT0YJCiD7XLw&e=>
> from default-token-2mz7g (ro)
> Volumes:
>   robot-logs:
>     Type:       PersistentVolumeClaim (a reference to a
> PersistentVolumeClaim in the same namespace)
>     ClaimName:  dev-robot-robot
>     ReadOnly:   false
>   dshm:
>     Type:       EmptyDir (a temporary directory that shares a pod's
> lifetime)
>     Medium:     Memory
>     SizeLimit:  <unset>
>   localtime:
>     Type:          HostPath (bare host directory volume)
>     Path:          /etc/localtime
>     HostPathType:
>   robot-eteshare:
>     Type:      ConfigMap (a volume populated by a ConfigMap)
>     Name:      dev-robot-robot-eteshare-configmap
>     Optional:  false
>   default-token-2mz7g:
>     Type:        Secret (a volume populated by a Secret)
>     SecretName:  default-token-2mz7g
>     Optional:    false
> QoS Class:       BestEffort
> Node-Selectors:  <none>
> Tolerations:     node.kubernetes.io/not-ready:NoExecute
> <https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_not-2Dready-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=0b-LBWz7xZSxyHD54a1S970SRnyV92DCFb8OC5hGa6g&e=>
> for 300s
>                  node.kubernetes.io/unreachable:NoExecute
> <https://urldefense.proofpoint.com/v2/url?u=http-3A__node.kubernetes.io_unreachable-3ANoExecute&d=DwMFaQ&c=LFYZ-o9_HUMeMTSQicvjIg&r=e3d1ehx3DI5AoMgDmi2Fzw&m=67kLb4PSwSX9Fj04dF9MkMMJ0A2LrX9Ze0p6tuDLxo4&s=lUcFt1b-6535RBbsO7aF5SFumNsJvLXBjdo1eIXnxgA&e=>
> for 300s
> Events:
>   Type     Reason   Age    From               Message
>   ----     ------   ----   ----               -------
>
> * Warning  Evicted  4m26s  kubelet, dublin06  The node was low on
> resource: ephemeral-storage. Container robot was using 127421328Ki, which
> exceeds its request of 0.   Normal   Killing  4m24s  kubelet, dublin06
>  Killing container with id docker://robot:Need to kill Pod*
>
> ----------------------------------
>
>
>
> thanks,
>
> -RPMishra
>
>
>
>
>
> On Fri, Mar 20, 2020 at 9:42 PM Steve Stark <[email protected]> wrote:
>
> Great that helps. It looks like instantiation worked, the part that's
> failing is some custom heat stack validation after the instantiation.
>
> I was able to replicate the failure you're getting. Can you change the
> admin_tenant value for your identity service to the tenant name instead of
> the tenant ID and try again? I think in your case the tenant name = admin.
>
> This is in the mariadb pod -
>
> MariaDB [catalogdb]> update identity_services set admin_tenant="<tenant
> name>" where id="<identity service id>";
>
> Thanks
> Steven
>
> 
>

-=-=-=-=-=-=-=-=-=-=-=-
Links: You receive all messages sent to this group.

View/Reply Online (#20413): https://lists.onap.org/g/onap-discuss/message/20413
Mute This Topic: https://lists.onap.org/mt/72026138/21656
Group Owner: [email protected]
Unsubscribe: https://lists.onap.org/g/onap-discuss/unsub  
[[email protected]]
-=-=-=-=-=-=-=-=-=-=-=-

Reply via email to