list-objects-filter: add list_objects_filter__filter_oidset()

The existing filter entry point, list_objects_filter__filter_object(),
is built around the object-walk path: it expects traversal context and
provisional omit sets, and is meant to be called as objects are
visited during a walk. A caller that already has a set of OIDs in hand
and only wants to know which ones a filter would select has no usable
entry point into the filter API.

--drop-filtered is exactly such a caller: it collects promisor blobs
into an oidset and needs to know which of them exceed the filter
threshold, without performing an object walk.

Add a helper, list_objects_filter__filter_oidset(), that takes a set
of OIDs and populates an "omitted" set with those that would be
filtered out by the given filter options. Only blob:limit=N filters
are supported for now.

This helper does not actually reuse the existing filter machinery.
It reimplements the blob:limit size check directly. That machinery
is tied to the object-walk path and cannot easily be driven
from a plain oidset. A NEEDSWORK comment marks this so the helper can
later be refactored to reuse the real filter logic instead of
duplicating it.

OBJECT_INFO_SKIP_FETCH_OBJECT is passed when reading object info so
the helper never triggers a lazy fetch.

Mentored-by: Christian Couder <christian.couder@gmail.com>
Mentored-by: Siddharth Asthana <siddharthasthana31@gmail.com>
Signed-off-by: Siddharth Shrimali <r.siddharth.shrimali@gmail.com>
Signed-off-by: Junio C Hamano <gitster@pobox.com>
This commit is contained in:
Siddharth Shrimali
2026-08-06 16:51:57 +05:30
committed by Junio C Hamano
parent 56f410ea37
commit ba07f620d5
2 changed files with 61 additions and 0 deletions

View File

@@ -828,3 +828,48 @@ void list_objects_filter__free(struct filter *filter)
filter->free_fn(filter->filter_data);
free(filter);
}
/*
* NEEDSWORK: this reimplements the blob:limit size check rather than
* reusing the existing filter machinery in
* list_objects_filter__filter_object(). That machinery is currently
* tied to the object-walk path and cannot easily be driven from a
* plain oidset. It would be nice to refactor the filter code so this
* helper can reuse it instead of duplicating the size check.
*/
int list_objects_filter__filter_oidset(struct repository *r,
struct list_objects_filter_options *opts,
const struct oidset *in,
struct oidset *omitted)
{
struct oidset_iter iter;
const struct object_id *oid;
if (opts->choice != LOFC_BLOB_LIMIT)
return error(_("filter_oidset: only blob:limit filters are supported"));
oidset_iter_init(in, &iter);
while ((oid = oidset_iter_next(&iter))) {
struct object_info info = OBJECT_INFO_INIT;
enum object_type type;
unsigned long size;
info.typep = &type;
info.sizep = &size;
/*
* Use OBJECT_INFO_SKIP_FETCH_OBJECT to avoid triggering
* a lazy fetch while inspecting candidates for removal.
*/
if (odb_read_object_info_extended(r->objects, oid, &info,
OBJECT_INFO_SKIP_FETCH_OBJECT) < 0)
continue;
if (type != OBJ_BLOB)
continue;
if (size >= opts->blob_limit_value)
oidset_insert(omitted, oid);
}
return 0;
}

View File

@@ -94,4 +94,20 @@ enum list_objects_filter_result list_objects_filter__filter_object(
*/
void list_objects_filter__free(struct filter *filter);
/*
* Given a set of OIDs in 'in', populate 'omitted' with those that
* would be filtered by 'opts'. Currently only blob:limit=N is
* supported. Objects that cannot be read are silently skipped.
*
* NEEDSWORK: this reimplements the blob:limit size check rather than
* reusing the existing filter machinery. See the matching comment in
* list-objects-filter.c.
*
* Return 0 on success, -1 if the filter is not supported.
*/
int list_objects_filter__filter_oidset(struct repository *r,
struct list_objects_filter_options *opts,
const struct oidset *in,
struct oidset *omitted);
#endif /* LIST_OBJECTS_FILTER_H */