Vladislav Glinskiy created SPARK-33841:
------------------------------------------

             Summary: Jobs disappear intermittently from the SHS under high load
                 Key: SPARK-33841
                 URL: https://issues.apache.org/jira/browse/SPARK-33841
             Project: Spark
          Issue Type: Task
          Components: Spark Core
    Affects Versions: 3.0.1, 3.0.0
         Environment: SHS is running locally on Ubuntu 19.04

conf/spark-defaults.conf:
{code:java}
spark.history.fs.logDirectory  s3a\://shs-reproduce-bucket/eventlog
spark.hadoop.fs.s3a.impl       org.apache.hadoop.fs.s3a.S3AFileSystem
spark.hadoop.fs.s3a.access.key <ACCESS-KEY>
spark.hadoop.fs.s3a.secret.key <SECRET-KEY>
{code}
 
            Reporter: Vladislav Glinskiy


Ran into an issue when a particular job was displayed in the SHS and 
disappeared after some time, but then, in several minutes showed up again.

The issue is caused by SPARK-29043, which is designated to improve the 
concurrent performance of the History Server. The 
[change|https://github.com/apache/spark/pull/25797/files#] breaks the ["app 
deletion" 
logic|https://github.com/apache/spark/pull/25797/files#diff-128a6af0d78f4a6180774faedb335d6168dfc4defff58f5aa3021fc1bd767bc0R563]
 because of missing proper synchronization for {{processing}} event log 
entries. Since SHS now [filters 
out|https://github.com/apache/spark/pull/25797/files#diff-128a6af0d78f4a6180774faedb335d6168dfc4defff58f5aa3021fc1bd767bc0R462]
 all {{processing}} event log entries, such entries do not have a chance to be 
[updated with the new 
{{lastProcessed}}|https://github.com/apache/spark/pull/25797/files#diff-128a6af0d78f4a6180774faedb335d6168dfc4defff58f5aa3021fc1bd767bc0R472]
 time and thus any entity that completes processing right after 
[filtering|https://github.com/apache/spark/pull/25797/files#diff-128a6af0d78f4a6180774faedb335d6168dfc4defff58f5aa3021fc1bd767bc0R462]
 and before [the check for stale 
entities|https://github.com/apache/spark/pull/25797/files#diff-128a6af0d78f4a6180774faedb335d6168dfc4defff58f5aa3021fc1bd767bc0R560]
 will be identified as stale and will be deleted from the UI until the next 
{{checkForLogs}} run. This is because [updated {{lastProcessed}} time is used 
as 
criteria|https://github.com/apache/spark/pull/25797/files#diff-128a6af0d78f4a6180774faedb335d6168dfc4defff58f5aa3021fc1bd767bc0R557],
 and event log entries that missed to be updated with a new time, will match 
that criteria.

The issue can be reproduced by generating a big number of event logs and 
uploading them to the SHS event log directory on S3. Essentially, around 
800(82.6 MB) copies of an event log file were created using 
[shs-monitor|https://github.com/vladhlinsky/shs-monitor] script. Strange 
behavior of SHS counting the total number of applications was noticed - at 
first, the number was increasing as expected, but with the next page refresh, 
the total number of applications decreased. No errors were logged by SHS.



--
This message was sent by Atlassian Jira
(v8.3.4#803005)

---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]

Reply via email to