On 09/22/2013 09:27 PM, Walter Bright wrote:
On 9/22/2013 11:43 AM, Timon Gehr wrote:
Tracking line numbers is likely worth it. I don't believe that
providing column
numbers in error messages necessitates a slowdown though.

Please consider that:

      IT ISN'T JUST FOR ERROR MESSAGES

It would go in the symbolic debug info, too, where it will be required
everywhere and will be right there on the fast path through the
lexer/compiler.
...

There is no such thing as a law that obliges compiler writers to add column numbers in debug info when such information is available in frontend error messages. The trade-offs involved in both cases may be different and deserve separate consideration.

Now consider the lexer doing a fast skip over comment text (this ranks
fairly high in the profile). This operation gets a lot slower if you're
also keeping track of column number.

I am not keeping track of column number.

Please note that:

      COLUMN NUMBER ISN'T THE OFFSET FROM THE START OF THE LINE
...

Obviously. I compute the correct column number exactly in the case when an error message should actually be printed. It is not necessary to do any of this on the fast path. The additional memory word per location that I waste in comparison to DMD could be shaved off by using more computation in the error case (or by giving up support for exact underlining), but the project has not yet reached a stage where this is worth considering/measuring.

Excerpts from actual code I wrote roughly two years ago:

class Source{
    // computes a slice of the entire first line
    // where some given slice occurs in the source buffer.
    // this allows to recover column information on the fly, and we
    // will also be able to print the line where an error occurred
    // without storing it explicitly.
    // running time is linear in output length
    string getLineOf(string rep)in{/*...*/}out{/*...*/}body{
        string before=code[0..rep.ptr-code.ptr];
        string after=code[rep.ptr-code.ptr..$];
        immutable(char)* start=code.ptr, end=code.ptr+code.length;

        // It is fine to skip decoding here, because we are just
        // searching for ASCII characters.
        // TODO: support unicode line breaks?
        foreach_reverse(ref c; before)
            if(c=='\n'||c=='\r'){start = &c+1; break;}
        foreach(ref c; after)
            if(c=='\n'||c=='\r'){end = &c; break;}
        return start[0..end-start];
    }
    // ...
}

struct Location{
    string rep;    // slice of the code representing the Location
    int line;      // line number at start of location

    @property Source source()const{
        auto src = Source.get(rep); // (currently just a linear search)
        assert(src, "source for '"~rep~"' not found!");
        return src;
    }
    // ...
}

int getColumn(Location loc, int tabsize){
    int res=1;
    auto l=loc.source.getLineOf(loc.rep);
    for(;!l.empty&&l[0]&&l.ptr<loc.rep.ptr; l.popFront()){
        if(l.front=='\t') res=res-res%tabsize+tabsize;
        else res++;
    }
    return res;
}



Reply via email to