In the first elostirion post I described the loop I wanted; to declare the state of a repository fleet, find the drift, show the fix, and open the pull requests that converge it.

The awkward word in that sentence was fleet.

The first version worked well when all the repositories were already on disk. That is useful for a demo and inconvenient for the exact situation elostirion is meant to solve. A fleet might be 20 repositories or > 200. They may live on GitHub or Bitbucket, use different default branches, and never share a parent directory on one machine. “Clone everything first” is not fleet management; it is setup work wearing a small hat.

Over the last two releases I have been removing that assumption. elo can now enumerate an organization, inspect repositories directly through the provider API, load the fleet spec remotely, and produce reports that are much safer to consume in both terminals and CI.

The command I wanted is now real:

elo scan --org github.com/acme \
  --spec https://github.com/acme/platform/blob/main/fleet-spec.yaml

No checkout directory. No loop around git clone. One spec, one fleet report.

one scanner, several filesystems⌗

The tempting implementation was to add GitHub calls inside every scanner. The Go module scanner would know how to fetch go.mod; the Python scanner would know how to fetch pyproject.toml; the Dockerfile scanner would know about the Contents API. That would work quickly, but age badly.

elo’s scanners already consumed Go’s fs.FS interface. I kept that boundary and built provider readers behind it instead:

type RemoteReader interface {
    GetFile(ctx context.Context, path string) ([]byte, error)
    ListFiles(ctx context.Context, dir string) ([]DirEntry, error)
}

An adapter exposes a RemoteReader as an fs.FS. From a scanner’s point of view, opening go.mod is the same operation whether the bytes came from a local checkout, GitHub’s Contents API, or Bitbucket’s source API.

That choice did more than avoid duplicated HTTP code. It preserved the rule that scanners extract facts and provider clients transport bytes. Authentication, pagination, refs, and provider-specific errors stay in pkg/reader; parsing and source locations stay in pkg/scan.

The payoff is visible in the feature itself. Remote scanning did not require a second set of Go, Python, and Dockerfile scanners. The existing scanners ran unchanged against a new filesystem.

discovering 500 repositories without guessing⌗

For organization scans, elostirion asks GitHub for repositories in pages of 100 and records each repository’s default branch. A 500-repository org takes five enumeration requests rather than one request per repository just to learn what exists.

Each repository is then root-listed before scanning. That serves two purposes: it fails early on authentication problems, and it lets language filters skip irrelevant repositories before parsing their contents. With --language go, a repository without go.mod does not run the Go scanner.

This is a meaningful reduction in wasted work, but it is not the end of the request-budget story. The current Contents API reader still performs file-level requests after discovery. Candidate paths, language detection, and scanners can ask for the same file more than once. At 500 repositories, small per-repository inefficiencies become thousands of calls and run directly into GitHub’s rate limit.

The next optimization is therefore measurable: cache reads within a repository and replace repeated Contents calls with a tree or batched query. I want to publish the before/after request count with the benchmark harness, not a number that merely sounds good. Fleet tooling should account for its API budget as carefully as it accounts for CPU or memory.

the spec became remote too⌗

Remote repositories exposed another local assumption: the fleet could be remote while its desired state still had to be a local file.

--spec can now resolve a remote GitHub location as well as a filesystem path. That matters beyond convenience. A centrally reviewed spec can live beside the platform code, and CI jobs can consume the same committed version without copying it into every repository.

It also creates a cleaner operating model:

  1. Review policy changes in one repository.
  2. Point scans and CI gates at that policy.
  3. Let elostirion report which repositories digress.
  4. Use plan and apply to converge them.

The desired state is versioned, reviewable, and independent of the machine running the scan.

reports are an API, even when they look like output⌗

The least glamorous work this week may be the work I trust most.

elo supports text, JSON, JUnit, and SARIF output. Text is for a person; the other three are contracts with scripts, test viewers, and code-scanning interfaces. A tiny rendering change can be a breaking change even when the terminal output still looks fine.

I added golden tests for every report format and immediately found the kind of failure golden tests are meant to expose: the fixtures were under testdata, but an ignore rule kept them out of Git. The tests passed on the machine that had generated the files and failed when CI received a clean checkout.

The fix was small. The lesson was larger: tests do not protect an artifact that never enters the repository.

While tightening the renderers, I also normalized empty findings in JSON and made empty expected/actual values explicit in JUnit and SARIF. These are minor details until another system parses them. Then they are the interface.

terminal output should know where it is running⌗

Color used to be an implicit terminal decision. That breaks down as soon as output is redirected, captured in CI, or compared in a golden test.

elo now exposes the familiar contract:

elo scan --color auto
elo scan --color always
elo scan --color never

auto follows terminal capability, always is useful when the consumer understands ANSI sequences, and never keeps logs and machine-adjacent output clean. It is a small flag, but it removes an entire category of “works in my terminal” behavior.

a Dockerfile has more than one base image⌗

The Dockerfile scanner originally recorded the final image. That is enough for a single-stage build and incomplete for the Go services I actually care about. A multi-stage Dockerfile often uses a Go image to compile and a distroless or Alpine image to run. A policy on the runtime image says nothing about whether the compiler still contains a vulnerable toolchain.

The scanner now records dockerfile.builder_image separately. It resolves a stage explicitly named builder, falling back to the first FROM, while continuing to track the final runtime image.

That makes policies precise:

- id: current-go-builder
  field: dockerfile.builder_image
  op: matches
  value: '^golang:1\.25'

One Dockerfile can now answer two different questions: what built this binary, and what will run it?

local fleets got deeper⌗

The local scanner had the opposite topology problem. It only discovered repositories one directory below the supplied root, so structures such as services/payments/api could disappear from a fleet scan.

Discovery is now recursive and bounded by --depth, with hidden directories, vendor, and node_modules pruned. The default depth is three: deep enough for common grouped repository layouts, finite enough to avoid turning a policy scan into an accidental filesystem crawl.

The bound is an important part of the feature. Recursive discovery without a budget is just an unbounded walk with optimistic branding.

profiling before folklore⌗

I also added an opt-in CPU profile and a PGO profile to the CLI build. The goal is not to claim performance wins from the existence of a .pgo file. It is to make performance work evidence-driven as organization scans become larger.

The next bottleneck is likely to be network shape rather than local parsing. Profiling lets me verify that instead of optimizing whichever function happens to look busy in the source.

what shipped⌗

Since the last post, the repository has moved through v0.2.0 & 0.3.0 with a few additional changes now on main:

  • clone-free scans of individual GitHub and Bitbucket repositories;
  • GitHub organization enumeration with pagination;
  • remote fleet-spec loading;
  • golden coverage for text, JSON, JUnit, and SARIF reports;
  • explicit --color always|auto|never behavior;
  • separate Docker builder-image facts;
  • bounded recursive repository discovery through --depth;
  • CPU profiling support and a PGO-enabled CLI build (not that 2% performance improvement makes a huge difference).

what I am working on next⌗

I am ’teaching’ elostirion to read pipeline policy across GitHub Actions, Bitbucket Pipelines, and GitLab CI. The next post will cover ordered steps, images, required files, and why pipeline YAML forced the fact model to grow.

elo is still the same bet: a repository policy should be declarative, inspectable, and able to close the loop without requiring a platform service. The difference is that it is starting to operate on a fleet as a fleet, rather than as a convenient pile of local directories.

The code is at github.com/whoisnjoguu/elostirion.

@CtrlAltEng on X

@CtrlAltEng — come say hi.