mirror of
https://github.com/git/git.git
synced 2026-08-08 17:11:48 +00:00
Blame and "git log --stat" recover hunk coordinates by diffing blob pairs, and recompute them on every run. Add a cache of those coordinates at $GIT_DIR/objects/info/diff-hunks, beside the commit-graph, so a later run can look them up instead of decompressing the blobs and running xdiff again. The store is a single chunk-format file (see gitformat-chunk(5)): an 8-byte header, a DHIX index of fixed-size entries sorted by key, a DHDT segment of hunk records, and a trailing hash checksum. An entry is keyed by the two blob object ids and the xdl_opts the pair was diffed under, so a stored result is served only where that exact key recurs, independent of path. A zero-context diff trims unchanged lines from hunk edges and can pick a different but equally valid set of hunks than an untrimmed diff, so a recording caller stores a pair only when its trimmed and untrimmed diffs are identical; such an entry answers any consumer at any context, and the rare divergent pair is always computed. Identical hunk blocks are interned once and shared across keys. The library provides a reader (repo_diff_hunks_store and _replay, gated by core.diffHunks), loaded once and cached on the object database as the commit-graph is, and a writer that accumulates entries and flushes them in one atomic pass. An absent, corrupt, or disabled store reads as all misses. A record with no hunks is invalid too: replaying it would claim the pair equivalent, which the store never asserts, so it reads as a miss. Ordinary reads are diagnostic-free. Loading parses the chunk table through read_table_of_contents_quiet(), new in chunk-format, which prints nothing on a malformed table and takes the repository's hash algorithm rather than the_hash_algo, so the file is bounds-checked under the algorithm it is keyed by. The flush closes the repository's mmapped store and forgets that loading was attempted before committing the lockfile. A warming run that also reads may hold the file it is replacing mapped, and the rename must not land on a live mapping, which Windows refuses; a read after the flush then observes the committed file. commit-graph closes its graph before committing for the same reason. Writing is off by default, enabled per run by GIT_DIFF_HUNKS_WRITE or persistently by diffHunks.write, the environment winning. A writer seeds from the existing store, so a flush merges rather than replaces. The seed's checksum is verified first: a corrupt store is discarded, not rewritten with a fresh checksum verify could no longer catch. An entry that fails the shared diff_provider_check_hunk() or names no blob is dropped with a warning, since it would only ever read as a miss. A seed that discarded or dropped anything forces the flush even when the warming run computed nothing new. The writer fsyncs through a new diff-hunks core.fsync component. "git diff-hunks" inspects and manages the file: "verify" checks the checksum, chunk table, sort order, entry bounds, and every entry's hunk sequence against that shared check, so a store whose entries could only read as misses fails verify; "clear" removes the file. Later patches wire the readers and the writer into the diff and blame paths. Signed-off-by: Michael Montalbo <mmontalbo@gmail.com> Signed-off-by: Junio C Hamano <gitster@pobox.com>
88 lines
2.6 KiB
C
88 lines
2.6 KiB
C
#ifndef CHUNK_FORMAT_H
|
|
#define CHUNK_FORMAT_H
|
|
|
|
#include "hash.h"
|
|
|
|
struct hashfile;
|
|
struct chunkfile;
|
|
|
|
#define CHUNK_TOC_ENTRY_SIZE (sizeof(uint32_t) + sizeof(uint64_t))
|
|
|
|
/*
|
|
* Initialize a 'struct chunkfile' for writing _or_ reading a file
|
|
* with the chunk format.
|
|
*
|
|
* If writing a file, supply a non-NULL 'struct hashfile *' that will
|
|
* be used to write.
|
|
*
|
|
* If reading a file, use a NULL 'struct hashfile *' and then call
|
|
* read_table_of_contents(). Supply the memory-mapped data to the
|
|
* pair_chunk() or read_chunk() methods, as appropriate.
|
|
*
|
|
* DO NOT MIX THESE MODES. Use different 'struct chunkfile' instances
|
|
* for reading and writing.
|
|
*/
|
|
struct chunkfile *init_chunkfile(struct hashfile *f);
|
|
void free_chunkfile(struct chunkfile *cf);
|
|
int get_num_chunks(struct chunkfile *cf);
|
|
typedef int (*chunk_write_fn)(struct hashfile *f, void *data);
|
|
void add_chunk(struct chunkfile *cf,
|
|
uint32_t id,
|
|
size_t size,
|
|
chunk_write_fn fn);
|
|
int write_chunkfile(struct chunkfile *cf, void *data);
|
|
|
|
int read_table_of_contents(struct chunkfile *cf,
|
|
const unsigned char *mfile,
|
|
size_t mfile_size,
|
|
uint64_t toc_offset,
|
|
int toc_length,
|
|
unsigned expected_alignment);
|
|
|
|
/*
|
|
* Like read_table_of_contents(), for a reader that treats a malformed
|
|
* table as an absent file rather than reporting it: nothing is printed
|
|
* on failure, and the trailing-checksum bound is computed with the
|
|
* given hash algorithm instead of the_hash_algo.
|
|
*/
|
|
int read_table_of_contents_quiet(struct chunkfile *cf,
|
|
const unsigned char *mfile,
|
|
size_t mfile_size,
|
|
uint64_t toc_offset,
|
|
int toc_length,
|
|
unsigned expected_alignment,
|
|
const struct git_hash_algo *algo);
|
|
|
|
#define CHUNK_NOT_FOUND (-2)
|
|
|
|
/*
|
|
* Find 'chunk_id' in the given chunkfile and assign the
|
|
* given pointer to the position in the mmap'd file where
|
|
* that chunk begins. Likewise the "size" parameter is filled
|
|
* with the size of the chunk.
|
|
*
|
|
* Returns CHUNK_NOT_FOUND if the chunk does not exist.
|
|
*/
|
|
int pair_chunk(struct chunkfile *cf,
|
|
uint32_t chunk_id,
|
|
const unsigned char **p,
|
|
size_t *size);
|
|
|
|
typedef int (*chunk_read_fn)(const unsigned char *chunk_start,
|
|
size_t chunk_size, void *data);
|
|
/*
|
|
* Find 'chunk_id' in the given chunkfile and call the
|
|
* given chunk_read_fn method with the information for
|
|
* that chunk.
|
|
*
|
|
* Returns CHUNK_NOT_FOUND if the chunk does not exist.
|
|
*/
|
|
int read_chunk(struct chunkfile *cf,
|
|
uint32_t chunk_id,
|
|
chunk_read_fn fn,
|
|
void *data);
|
|
|
|
uint8_t oid_version(const struct git_hash_algo *algop);
|
|
|
|
#endif
|