mercurial: tests/test-bdiff.py.out@9a8363d23419 (annotated)

bdiff: deal better with duplicate lines The longest_match code compares all the possible positions in two files to find the best match. Given a pair of sequences, it effectively searches a grid like this: a b b b c . d e . f 0 1 2 3 4 5 6 7 8 9 a 1 - - - - - - - - - b - 2 1 1 - - - - - - b - 1 3 2 - - - - - - b - 1 2 4 - - - - - - . - - - - - 1 - - 1 - Here, the 4 in the middle says "the first four lines of the file match", which it can compute be comparing the fourth lines and then adding one to the result found when comparing the third lines in the entry to the upper left. We generally avoid the quadratic worst case by only looking at lines that match, which is precomputed. We also avoid quadratic storage by only keeping a single column vector and then keeping track of the best match. Unfortunately, this can get us into trouble with the sequences above. Because we want to reuse the '3' value when calculating the '4', we need to be careful not to overwrite it with the '2' we calculate immediately before. If we scan left to right, top to bottom, we're going to have a problem: we'll overwrite our 3 before we use it and calculate a suboptimal best match. To address this, we can either keep two column vectors and swap between them (which significantly complicates bookkeeping), or change our scanning order. If we instead scan from left to right, bottom to top, we'll avoid ever overwriting values we'll need in the future. This unfortunately needs several changes to be made simultaneously: - change the order we build the initial hash chains for the b sequence - change the sentinel values from INT_MAX to -1 - change the visit order in the longest_match inner loop - add a tie-breaker preference for earlier matches This last is needed because we previously had an implicit tie-breaker from our visitation order that our test suite relies on. Later matches can also trigger a bug in the normalization code in diff().

400 8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	1	*** 'a\nc\n\n\n\n' 'a\nb\n\n\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	2	*** 'a\nb\nc\n' 'a\nc\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	3	*** '' ''
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	4	*** 'a\nb\nc' 'a\nb\nc'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	5	*** 'a\nb\nc\nd\n' 'a\nd\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	6	*** 'a\nb\nc\nd\n' 'a\nc\ne\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	7	*** 'a\nb\nc\n' 'a\nc\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	8	*** 'a\n' 'c\na\nb\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	9	*** 'a\n' ''
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	10	*** 'a\n' 'b\nc\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	11	*** 'a\n' 'c\na\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	12	*** '' 'adjfkjdjksdhfksj'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	13	*** '' 'ab'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	14	*** '' 'abc'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	15	*** 'a' 'a'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	16	*** 'ab' 'ab'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	17	*** 'abc' 'abc'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	18	*** 'a\n' 'a\n'
8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	19	*** 'a\nb' 'a\nb'
7104 9514cbb6e4f6 bdiff: normalize the diff (issue1295) Benoit Boissinot <benoit.boissinot@ens-lyon.org> parents: 400 diff changeset	20	6 6 'y\n\n'
9514cbb6e4f6 bdiff: normalize the diff (issue1295) Benoit Boissinot <benoit.boissinot@ens-lyon.org> parents: 400 diff changeset	21	6 6 'y\n\n'
9514cbb6e4f6 bdiff: normalize the diff (issue1295) Benoit Boissinot <benoit.boissinot@ens-lyon.org> parents: 400 diff changeset	22	9 9 'y\n\n'
29013 9a8363d23419 bdiff: deal better with duplicate lines Matt Mackall <mpm@selenic.com> parents: 15530 diff changeset	23	0 0 'a\nb\nb\n'
9a8363d23419 bdiff: deal better with duplicate lines Matt Mackall <mpm@selenic.com> parents: 15530 diff changeset	24	12 12 'b\nc\n.\n'
9a8363d23419 bdiff: deal better with duplicate lines Matt Mackall <mpm@selenic.com> parents: 15530 diff changeset	25	16 18 ''
400 8b067bde6679 Add a fast binary diff extension (not yet used) mpm@selenic.com parents: diff changeset	26	done
15530 eeac5e179243 mdiff: replace wscleanup() regexps with C loops Patrick Mezard <pmezard@gmail.com> parents: 8449 diff changeset	27	done

author	Matt Mackall <mpm@selenic.com>
	Thu, 21 Apr 2016 21:05:26 -0500
branch	stable
changeset 29013	9a8363d23419
parent 15530	eeac5e179243
child 30427	ede7bc45bf0a
permissions	-rw-r--r--