dpol1 commented on code in PR #2181:
URL: https://github.com/apache/stormcrawler/pull/2181#discussion_r4104400887


##########
external/ai/src/main/java/org/apache/stormcrawler/ai/AbstractLLMTextExtractor.java:
##########
@@ -51,11 +57,20 @@ public abstract class AbstractLLMTextExtractor implements 
TextExtractor {
     public static final String USER_PROMPT = "textextractor.llm.prompt";
     public static final String USER_REQUEST = "textextractor.llm.user_request";
     public static final String LISTENER_CLASS = 
"textextractor.llm.listener.clazz";
+    public static final String TEXT_MAX_LENGTH = 
"textextractor.llm.text.maxlength";

Review Comment:
   we already have `textextractor.skip.after` on `TextExtractor` for this, can 
we reuse it? the README line saying it's unsupported here can then go



##########
external/ai/src/main/java/org/apache/stormcrawler/ai/AbstractLLMTextExtractor.java:
##########
@@ -149,15 +170,84 @@ public String text(Object element) {
 
     /**
      * Replaces placeholders in the user message template with the actual HTML 
content and user
-     * request.
+     * request. Marker tokens of the template, such as {@code 
<|HTML_CONTENT_END|>}, are removed
+     * from the HTML first so that the page cannot close or open a section of 
the prompt.
      *
      * @param userMessage the original user message template
      * @param html the HTML string to insert
      * @return the updated user message string with placeholders replaced
      */
     protected String replacePlaceholders(String userMessage, String html) {
-        userMessage = userMessage.replace("{HTML}", html);
+        // the request is substituted first so that a {REQUEST} in the page is 
left as it is
         userMessage = userMessage.replace("{REQUEST}", userRequest);
-        return userMessage;
+        return userMessage.replace("{HTML}", removeMarkers(html));
+    }
+
+    /**
+     * Returns the text of a reply: the content of its {@code <content>} 
envelope if there is one,
+     * or the whole reply otherwise, without markup and truncated to the 
length set with {@value
+     * #TEXT_MAX_LENGTH}.
+     *
+     * @param reply the text returned by the model
+     * @return the extracted text
+     */
+    protected String cleanReply(String reply) {
+        if (reply == null) {
+            return "";
+        }
+        final int start = reply.indexOf(CONTENT_START);
+        if (start >= 0) {
+            final int end = reply.lastIndexOf(CONTENT_END);
+            final int from = start + CONTENT_START.length();
+            reply = end >= from ? reply.substring(from, end) : 
reply.substring(from);
+        }
+        String text = stripMarkup(reply).strip();
+        if (textMaxLength >= 0 && text.length() > textMaxLength) {
+            int cut = textMaxLength;
+            if (cut > 0 && Character.isHighSurrogate(text.charAt(cut - 1))) {
+                cut--;
+            }
+            text = text.substring(0, cut);
+        }
+        return text;
+    }
+
+    /** keeps the text nodes of the input and their line breaks, dropping 
elements and comments */
+    private static String stripMarkup(String text) {

Review Comment:
   this eats the code blocks the prompt asks for: `List<String>` comes back as 
`List`, `x<y and y>z` as `xz`. strip only outside the ``` fences, or keep the 
envelope and skip the stripping?



##########
external/ai/src/main/java/org/apache/stormcrawler/ai/AbstractLLMTextExtractor.java:
##########
@@ -149,15 +170,84 @@ public String text(Object element) {
 
     /**
      * Replaces placeholders in the user message template with the actual HTML 
content and user
-     * request.
+     * request. Marker tokens of the template, such as {@code 
<|HTML_CONTENT_END|>}, are removed
+     * from the HTML first so that the page cannot close or open a section of 
the prompt.
      *
      * @param userMessage the original user message template
      * @param html the HTML string to insert
      * @return the updated user message string with placeholders replaced
      */
     protected String replacePlaceholders(String userMessage, String html) {
-        userMessage = userMessage.replace("{HTML}", html);
+        // the request is substituted first so that a {REQUEST} in the page is 
left as it is
         userMessage = userMessage.replace("{REQUEST}", userRequest);
-        return userMessage;
+        return userMessage.replace("{HTML}", removeMarkers(html));
+    }
+
+    /**
+     * Returns the text of a reply: the content of its {@code <content>} 
envelope if there is one,
+     * or the whole reply otherwise, without markup and truncated to the 
length set with {@value
+     * #TEXT_MAX_LENGTH}.
+     *
+     * @param reply the text returned by the model
+     * @return the extracted text
+     */
+    protected String cleanReply(String reply) {
+        if (reply == null) {
+            return "";
+        }
+        final int start = reply.indexOf(CONTENT_START);
+        if (start >= 0) {
+            final int end = reply.lastIndexOf(CONTENT_END);
+            final int from = start + CONTENT_START.length();
+            reply = end >= from ? reply.substring(from, end) : 
reply.substring(from);
+        }
+        String text = stripMarkup(reply).strip();
+        if (textMaxLength >= 0 && text.length() > textMaxLength) {
+            int cut = textMaxLength;
+            if (cut > 0 && Character.isHighSurrogate(text.charAt(cut - 1))) {
+                cut--;
+            }
+            text = text.substring(0, cut);
+        }
+        return text;
+    }
+
+    /** keeps the text nodes of the input and their line breaks, dropping 
elements and comments */
+    private static String stripMarkup(String text) {
+        final StringBuilder sb = new StringBuilder(text.length());
+        Parser.htmlParser()
+                .parseInput(text, "")
+                .body()
+                .traverse(
+                        (node, depth) -> {
+                            if (node instanceof TextNode t) {
+                                sb.append(t.getWholeText());
+                            }
+                        });
+        return sb.toString();
+    }
+
+    private String removeMarkers(String html) {

Review Comment:
   wouldn't `html.replace("<|", "< |")` do the same in one pass, whatever the 
template? no marker set, no loop, and nested markers stay linear



-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to