OK, so 100 is the cutoff whether to use a hash table or not. It is still odd that it only affects a relatively short block at the beginning of the repeated targets, though.
> On 16 Jun 2026, at 11.54, Peter Dalgaard <[email protected]> wrote: > > Something definitely looks odd... I certain can't think of a reason why the > behaviour would be keyed to the _size_ of the first argument like this: > >> pmatch(as.character(x)[1:100], as.character(x)) > [1] 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 > [19] 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 > [37] 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 > [55] 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 > [73] 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 > [91] 91 92 93 94 95 96 97 98 99 100 >> pmatch(as.character(x)[1:101], as.character(x)) > [1] 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 > [19] 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 > [37] 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 NA NA NA > [55] NA NA 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 > [73] 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 > [91] 91 92 93 94 95 96 97 98 99 100 101 >> pmatch(as.character(x)[2:102], as.character(x)) > [1] 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 > [19] 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 > [37] 38 39 40 41 42 43 44 45 46 47 48 49 50 51 1 NA NA NA > [55] NA 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 > [73] 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 > [91] 92 93 94 95 96 97 98 99 100 101 102 >>> pmatch(as.character(x)[3:102], as.character(x)) > [1] 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 > [19] 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 > [37] 39 40 41 42 43 44 45 46 47 48 49 50 51 1 2 54 55 56 > [55] 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 > [73] 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 > [91] 93 94 95 96 97 98 99 100 101 102 > > It's happening in the C code, though, so some poking around is required. > > - pd > > >> On 15 Jun 2026, at 22.14, Duncan Murdoch <[email protected]> wrote: >> >> I think your example is overly complicated. Wouldn't it be enough to show >> one vector that gives a bad result? For example: >> >> x <- c(1:51, 1:51) >> pmatch(x, x) >> >> which gives >> >> [1] 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 >> 19 20 21 22 23 24 25 26 27 28 29 >> [30] 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 >> 48 49 50 51 NA NA NA NA NA 57 58 >> [59] 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 >> 77 78 79 80 81 82 83 84 85 86 87 >> [88] 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 >> >> (with the NAs showing up at locations 52 to 56). >> >> I'm not 100% sure that's a bug, since the documentation for pmatch doesn't >> discuss what should happen if table (the second argument) contains >> duplicates. I think I'd agree with you that you should get 1:102 as the >> output, but maybe that was never intended to be supported. >> >> Duncan Murdoch >> >> On 2026-06-15 3:12 p.m., François Rousset via R-devel wrote: >>> Dear R-devel list, >>> the following code shows NA's appearing in the result of pmatch() when >>> comparing a vector to itself when the length of the vector is more than 100. >>> The main specificity of this example is that elements are repeated in >>> the vector which is matched to itself. >>> The results when they do not include any NA (i.e. for argument of length >>> <=100) are exactly as I expect from the documentation. >>> I see that the source C code uses distinct algorithms whether n_input <= >>> 100 || n_target <= 100 or not. Could there by a problem in the source >>> code for larger values ? >>> The NA's correspond to the first positions of the second replicate of >>> the integer sequence, e.g. positions 52 to 56 if n=51 in the example below. >>> But as shown below, the NA's do not appear when the sequence is reversed >>> before being compared to itself. >>> This was first detected with R 4.5.3 and is reproducible with a >>> just-downloaded, virgin R-devel installation. >>> Thanks in advance for any feedback, >>> F. >>> ====================== >>> countNAs <- function(n, rev.=FALSE) { >>> seqn <- seq(n) >>> if (rev.) seqn <- rev(seqn) >>> seqn <- rep(seqn,2) # of length 2 n: NA's >>> appear when 2 n > 100 >>> chk <- pmatch(seqn, seqn) # I expect the result to be >>> seq(2*n) >>> sum(is.na(chk)) >>> } >>> sapply(1:100, countNAs ) >>> sapply(1:100, countNAs , rev.=TRUE) >>> ======================= >>> Results: >>>> sapply(1:100, countNAs ) >>> [1] 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 >>> 0 0 0 0 0 0 >>> [29] 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 >>> 5 5 5 5 5 5 >>> [57] 5 5 5 6 6 6 6 6 6 6 6 6 6 7 7 7 7 7 7 7 7 7 >>> 7 8 8 8 8 8 >>> [85] 8 8 8 8 8 9 9 9 9 9 9 9 9 9 9 10 >>>> sapply(1:100, countNAs , rev.=TRUE) >>> [1] 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 >>> 0 0 0 0 0 0 0 0 0 >>> [43] 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 >>> 0 0 0 0 0 0 0 0 0 >>> [85] 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 >>>> sessionInfo() >>> R Under development (unstable) (2026-06-14 r90150 ucrt) >>> Platform: x86_64-w64-mingw32/x64 >>> Running under: Windows 10 x64 (build 19045) >>> Matrix products: default >>> LAPACK version 3.12.1 >>> locale: >>> [1] LC_COLLATE=French_France.utf8 LC_CTYPE=French_France.utf8 >>> [3] LC_MONETARY=French_France.utf8 LC_NUMERIC=C >>> [5] LC_TIME=French_France.utf8 >>> time zone: Europe/Paris >>> tzcode source: internal >>> attached base packages: >>> [1] stats graphics grDevices utils datasets methods base >>> loaded via a namespace (and not attached): >>> [1] compiler_4.7.0 tools_4.7.0 >>> ______________________________________________ >>> [email protected] mailing list >>> https://stat.ethz.ch/mailman/listinfo/r-devel >> >> ______________________________________________ >> [email protected] mailing list >> https://stat.ethz.ch/mailman/listinfo/r-devel > > -- > Peter Dalgaard, Professor, > Center for Statistics, Copenhagen Business School > Solbjerg Plads 3, 2000 Frederiksberg, Denmark > Phone: (+45)38153501 > Office: A 4.23 > Email: [email protected] Priv: [email protected] > -- Peter Dalgaard, Professor, Center for Statistics, Copenhagen Business School Solbjerg Plads 3, 2000 Frederiksberg, Denmark Phone: (+45)38153501 Office: A 4.23 Email: [email protected] Priv: [email protected] ______________________________________________ [email protected] mailing list https://stat.ethz.ch/mailman/listinfo/r-devel
