Hi,

> Also UUID is a bit odd. It is said to be 128 bits (16 bytes = 2*bigint)
> long.
> So it could be more generically expressed as BINARY(16)

UUID actually is not just a 16 bytes but it's a type strictly defined
by RFC. The most important it has well defined string representation
to be easily interchangable between different kinds of software. So
higher level of generality you are talking about will lack this
feature.

> BINARY(bytes)
> VARBINARY(bytes)
>
> These are simple for everyone to understand because the same idea has
> been implemented for strings.

AFAIK in h2 the only difference between VARCHAR and CHAR that latter
is trimmed from right and actually is not recommended for use because
it can behave differently in other dbms. I think the same thing would
be with VARBINARY/BINARY (in h2 there is no difference at all). So
just use VARCHAR and VARBINARY (or may be BLOB) probably will be good
engineering practice:)

But I agree that good user defined type system would be nice.

regards,
S.Vladykin

On 14 окт, 19:19, Rami <[email protected]> wrote:
> > The original question ultimately boils down to whether the author
> > wants to keep H2 a "true & pure DBMS", or allow adding some goodies
> > that are handy in some situations.
>
> I don't think the issue is best captured by setting against each other
> some obscure ideology of
> truth and purity with the style of handy and dandy goodies.
>
> I think the issue is how clearly some specific program is defined.
> We all know that most of the programs we write is meant for one
> particular use case.
> And then we keep adding features as the need arises.
>
> But when we try to construct a software that is meant to be used for
> some specific purpose by all kinds of software,
> then defining features in a manner that is logical and coherent makes it
>
> 1) easier to implement because the implementation flows from well
> thought out concepts
> 2) more free from bugs because the lack of endless combinations of
> corner cases
> 3) simpler to document because the concepts have been well defined from
> the outset
> and because the lack of corner cases.
>
> The best of both worlds can be achieved if the tool/program has a clear
> extension mechanism.
> In case of dbms these mechanisms include user defined types, functions
> and triggers.
> Then all kinds of handy and dandy features could be shared by the community
> but would not give headache to the core definitions of the dbms.
>
> Currently from these three extension mechanisms H2 is lacking only a
> good model of user defined types
> (I have argued before that the implementation of domain is still rather
> weak).
>
> But before implementing User Defined Types (domains) it would be
> beneficial to have a clear
> definition of what type means in H2.
>
> Let me give an example:
>
> - Type defines a set of values.
> - A variable that has that type must have a single value from that set
> of values.
> - Type needs to define a literal representation for all those values so
> that a literal always picks
> out exactly one of the values from that set.
> - Type needs a storage representation that the normal dbms user does not
> need to know anything about
> but a type definer must define it - if only implicitly.
>
> Type can have operators defined.
> The most important out of these operators is the equality operator.
> If a type does not have one defined it can not be used in a key
> (uniqueness not defined).
> It also can not be used in expression like A = B.
>
> So even though the equality operator is not strictly needed without it
> the type is obviously quite crippled.
>
> Composite types (types that are composed of many parts that are supposed
> to be visible to the user)
> have one basic deficiency.
>
> Example:
> Let's define a type
> Point(INT x, INT y).
> Then let's have 2 variables of that type:
> a = Point(2, 4)
> b = Point(1, 3)
> The equality operator would be defined so that it returns true if a.x =
> b.x and a.y = b.y
> One would need to define a component extractor operator (here .x and .y)
> in order to make queries regarding x and y coordinates.
> That's all fine but it is just more work than putting the data into two
> int columns.
> However if the user's interest concerns properties of the point say
> distance from origin,
> then it is well-grounded to define type Point.
> WHERE DIST_FROM_ORIGIN(a) > DIST_FROM_ORIGIN(b)
>
>  From the point of these definitions I think that OTHER and ARRAY are
> not very well defined.
> Currently OTHER does not seem to differ much from BINARY type.
> What you were asking was that equality operator would be defined for it.
> I would quess that now the equality is determined the same way as for
> BINARY type - that is comparing all the bytes with each other.
>
> Now let's consider the implementation.
> How are binary type values searched from the database file?
> I don't know, but I quess that one needs to read them all into memory
> and start comparing.
> How could this be sped up with an index?
> Maybe putting the first N bytes in some order and using binary search on
> that index, then loading only those
> binary values that matched and finally throwing away those that did not
> match after all.
> Obviously this strategy would fail if files of certain kind always start
> with the same kind of byte sequence.
>
> If we would define the equality operator for the OTHER type based on
> it's equals(obj) method then how would we define the index?
> If we don't define the index then to search all equal values would
> require us not only to load the bytes
> into memory but also to instantiate all those objects.
>
> So clearly we would need that the object would also have a method
> someting like long getIndexValue()
> that would also define an ordering in the type so we could find equal
> objects a bit faster.
>
> All in all my point is that I am not advocating some fanaticism towards
> purity & truth.
> I am advocating good engineering practices.
>
> - Rami
>
> Ps. Also I would criticize IDENTITY type because it is not a type but a
> shorthand to define a type (BIGINT) and a sequence with one word.
> Also UUID is a bit odd. It is said to be 128 bits (16 bytes = 2*bigint)
> long.
> So it could be more generically expressed as BINARY(16) (also a
> distinction between BINARY and VARBINARY should be made here).
> Actually it is not clear from the documentation if 16 bytes is the upper
> limit or if all values are 16 bytes.
> Also it is not said how the missing bytes are padded (probably padding
> is added to left with 0 bits).
> Also when a String is converted to bytes if the jvm encoding is UTF-8 it
> could produce uneven number of bytes but maybe this is taken care of
> with padding.
> Equality is obviously based on binary comparison.
> But wouldn't it have been a more generic strategy to define 2 type
> generators
>
> BINARY(bytes)
> VARBINARY(bytes)
>
> These are simple for everyone to understand because the same idea has
> been implemented for strings.
> And what's nice with this strategy is that it is compatible with SQL and
> JDBC!!! :-)

-- 
You received this message because you are subscribed to the Google Groups "H2 
Database" group.
To post to this group, send email to [email protected].
To unsubscribe from this group, send email to 
[email protected].
For more options, visit this group at 
http://groups.google.com/group/h2-database?hl=en.

Reply via email to