Missatge de Jean Louis Faucher <[email protected]> del dia dl., 29 de
juny 2026 a les 19:31:

>
>
> On 22 Jun 2026, at 15:47, Josep Maria Blasco <[email protected]>
> wrote:
>
> Bonjour Jean Louis,
>
> It's impressive what you can do with what you have implemented,
> thank you very much for your work.
>
> I'm wondering about how we should abstract things.
> Currently, RexxUnicodeServices is a wrapper for utf8Proc.
> But utf8Proc is _an implementation_ of the Unicode services,
> and it is conceivable that a different implementation could be used
> in the future.
>
>
>
> +1
>
> The code organization is ready for several libraries.
>
> classes/
>     support/
>         Unicode/
>             utf8proc/
>             ziglyph/ (tested but not retained)
>             uni-algo/ (idem)
>             UnicodeServices.cpp
>             UnicodeServices.hpp
>
> In the past, I tested several libraries together, all exposed via the same
> ooRexx class.
>
>
>
> Maybe RexxUnicodeServices could offer abstract services,
> like . RexxUnicodeServices~utf8.toNFC,  utf8.toNFD, etc.,
> which would of course be currently implemented using utf8Proc.
> utf8Proc could also be  exposed, of course, maybe with its own class,
> but the recommended way to use low-level Unicode would be to use
> the abstract services class.
>
>
>
> IMO, the methods currently exposed by RexxUnicodeServices are already
> abstracted from the underlying Unicode library.
> The only visible dependency on utf8proc is utf8procVersion().
>

My impression is that the presence of utf8Transform is
misleading. utf8Transform is very utf8proc-specific, it is a Swiss-army
knife function which does a number of things. For example, you use it in
rxunicode.cls to implement normalization. But a user of RexxUnicodeServices
would expect to see abstracted pure Unicode functions, not the details of
utf8proc.


> Review of the methods currently exposed by RexxUnicodeServices:
>
> https://github.com/jlfaucher/executor5-bulk/blob/ec128b54ba8fd27933609dba9b2f422bebbf06d5/main/trunk/interpreter/classes/support/Unicode/UnicodeServices.hpp#L71-L100
>
>
> utf8DecodeCodepoint and utf8EncodeCodepoint are named "utf8" because they
> are specialized for UTF-8.
> Equivalent utf16, utf32, wtf8, etc. methods could be added (they are not
> supported by utf8proc).
>
>
> utf8Transform is an abstraction of several utf8proc functions.
> It is named "utf8" because it operates on UTF-8 strings.
> This method takes many arguments and would indeed impose this same set of
> arguments on other libraries.
>     use strict arg string, casefold = .false, lump= .false, nlf = 0,
> normalization = 0, stripCC = .false, stripIgnorable= .false, stripMark =
> .false, stripNA = .false
> This is the method used to implement NFC, NFD, etc.
> Some transformations, such as lump and nlf, are specific to utf8proc, but
> they could easily be supported by another library or simply rejected as
> unsupported.
> Given these eight options, and the possibility that other libraries may
> support more or fewer options, a natural abstraction would be to pass a
> table of options.
> I experimented with such a table, but quickly abandoned the idea.
> A good way to support so many options would be to use named arguments.
>
>
> The code point properties are returned in a form that is independent of
> utf8proc.
> For example:
>     RexxInteger *codepointCategory(RexxInteger *rexxCodepoint,
> VariableReference *refCode, VariableReference *refLabel); // enum
> refCode and refLabel are not a utf8proc feature.
> They are handled in UnicodeServices.cpp, which converts a utf8proc value
> into the corresponding Unicode Standard value.
>
>
> The lower/upper/title methods have no visible dependency on utf8proc.
> They take a code point and return either a code point or a boolean.
> They currently cover only the simple case mappings defined in
> UnicodeData.txt; they always return a single codepoint.
> Because of this limitation, I experimented with other libraries such as
> ziglyph and uni-algo.
> It was easy to add another method providing full support for the complex
> rules defined in SpecialCasing.txt.
> Currently, it is only a placeholder in rxunicode.cls.
>
>
> Additional methods could be added to RexxUnicodeServices.
> Or they could be added to a subclass of RexxUnicodeServices, for example a
> class named RexxUnicode.
> I think the distribution of methods between these two classes will be
> driven by the methods needed internally by the interpreter.
> Currentlly, we have one internal dependency on utf8proc for RFE #864 (at
> that time, RexxUnicodeServices did not yet exist).
>
>
> You can have a look at one possible extension of RexxUnicodeServices here:
>
> https://github.com/jlfaucher/executor5-bulk/blob/main/main/trunk/extensions/unicode/rxunicode.cls
> It provides only basic services, there is no BIF or methods for the String
> class.
> This file is located under trunk/unicode only because that simplifies
> packaging for the portable version.
>

> Examples (also used as regression tests):
>
> https://htmlpreview.github.io/?https://github.com/jlfaucher/executor5-bulk/blob/main/scripts/unicode/diary_examples-out.html
>
> That's impressive. Thanks for sharing it.

>
> The documentation is still in progress. I will create a second release
> once it is finalized.
>

Looking forward to read it.


> And perhaps it will be possible to include the new TUTOR with its strong
> integration into the interpreter (JMB WIP).
>

That would be nice :)

  Josep Maria


>
>
> _______________________________________________
> Oorexx-devel mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/oorexx-devel
>
_______________________________________________
Oorexx-devel mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/oorexx-devel

Reply via email to