arturobernalg commented on PR #890: URL: https://github.com/apache/httpcomponents-client/pull/890#issuecomment-6011653938
> Both SSE entity consumers break the event stream on LF only and just trim a trailing CR, so a lone CR (one not followed by LF) is kept inside the line instead of ending it; that diverges from the SSE line grammar, which terminates a line on CR, LF or CRLF, and mis-frames a stream that uses bare-CR separators. It also has a security edge: a server can place a lone CR in an id: value, which then survives into the parsed id and is copied straight into the Last-Event-ID header on the next reconnect, so the origin ends up controlling a carriage return in one of our own outgoing request headers. This makes a lone CR end the line in both the char and byte consumers while keeping CRLF a single break, so the id no longer carries a CR and bare-CR streams are framed correctly, with a regression test added to each consumer. Hi @dxbjavid I think there is still one edge case in `ByteSseEntityConsumer`. If the stream starts with a lone CR, that first byte is consumed while BOM detection is being resolved and is appended directly to `lineBuf`, bypassing the new CR/LF handling. This valid SSE input: ```java "\rdata: v\r\r" ``` should produce `data = "v"`, but currently produces `null`. I think the first non-BOM byte should go through the same byte-processing path as the normal parsing loop rather than being appended directly. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
