On Wed, Aug 20, 2008 at 5:13 PM, Roy Lyseng <[EMAIL PROTECTED]> wrote: > > > Jay Pipes wrote: >> >> Baron Schwartz wrote: >>> >>> If table meta-data is stored in tables, and InnoDB is the default >>> storage engine, is there any reason DDL should not be transactional? >> >> I don't see why metadata must be stored in a table. Sure, given a >> flexible API, it *could* be, but storing the metadata in a storage >> engine's tables is just about persistence. We could just as easily >> persist the metadata in a file using proto buffer's >> Message.write(CodedOutputStream *out) methods. > > If you use the storage engine to also commit DDL changes, commit handling of > DDL is simplified. With separate DDL persistence, you need a kind of 2PC. An > exception is DDL operations that deal with file creation and deletion... >> >> I think that something that is more important than transactional safety >> of DDL commands is having those commands be performed online, with no >> (or little) locking out of readers while the write operations are being >> performed. >> >> What enables online operations for DDL commands has to do with the way >> that metadata (I've started referring to this as catalog data) is >> managed in the kernel. If the catalog itself allows for versioning of >> catalog objects (e.g. a schema table definition), then a storage engine >> can rebuild or alter a schema table based on a newly supplied definition >> behind the scenes. When the rebuild is done, the catalog is notified of >> the update, locks the pointer to a catalog object and swaps the pointer >> with a newly updated one. Readers of the old version continue to see >> the old definition and may work on that definition and storage area. >> New readers will see the new definition and once no more readers of the >> old definition are around, the catalog can notify the storage engine >> that it can free up the resources that were dependent on the old >> definition. >> >> Of course, it's more complex than this, but this is the basic idea of >> online DDL operations... Proto buffers plays a central role in all of >> this. The generated classes from the .proto files will be passed to and >> from the storage engine to the catalog, effectively enabling online >> communication of changes in the schema table. > > It is a bit more complex because if you do an ALTER TABLE, you may actually > have to update two instances of the same table while the modification is > going on. If the ALTER is successful, the old table contents are scrapped > when all old references (transactions) are ended. If it is not, then the new > table contents are deleted, and everything is reverted. > > This is actually the strategy implemented in the Clustra/HADB product.
I think we can also potentially look to other products and how they do this. MVCC might enable transactional, non-blocking ALTER TABLE, for example. Instead of building a whole new table and copying each row into it, you could version the rows. Copy each row into a new row, and mark the old one as deleted in the ordinary MVCC way. Then an ALTER really just looks like an UPDATE. I realize I'm totally waving my hands about the "hard" parts of this. But it is doable. _______________________________________________ Mailing list: https://launchpad.net/~drizzle-discuss Post to : [email protected] Unsubscribe : https://launchpad.net/~drizzle-discuss More help : https://help.launchpad.net/ListHelp

