•
12 min read
The package name was not the network boundary
A Cortex API build failure that looked like flaky package downloads, and turned out to be a CUDA dependency graph, a 2 GiB egress ceiling, an HTTP proxy boundary, and a model distributed across several hosts.

The Cortex API image kept failing while pip downloaded a package it did not need.

The error was not particularly illuminating. pip reported an incomplete download and an SSLEOFError, which is a reasonable way for a client to describe a connection that stopped while bytes were still arriving. It looked like a network problem. The next run failed in the same general part of the install. The one after that did too.

That repeatability was the clue.

Cortex uses PyTorch for embeddings. Cortex does not require CUDA. Yet pip install -r requirements.txt was resolving torch==2.7.1 from ordinary PyPI, which means a Linux x86_64 install receives the CUDA-enabled distribution and its NVIDIA runtime dependencies. The build was downloading gigabytes of GPU software into an image that would use the CPU.

The connection was not randomly failing. Canopy Forest’s VM-isolated image-build path has a cumulative 2 GiB egress budget. Once the build had transferred that much through its proxy, the proxy closed the connection. pip happened to be the process holding the other end of it.

The immediate repair was one line. The more interesting repair was to stop treating a dependency name as a description of its network behaviour.

The build had a budget

The image build runs in a VM-isolated Forest job. It does not get the ordinary network of a long-lived machine. When a pipeline declares external access, the build receives a small egress proxy and an allowlist for that run. The proxy is the only external route available to the sandbox.

It also maintains a byte counter for the lifetime of the run. The current ceiling for a VM-isolated build is 2 GiB:

vmIsolationEgressMaxBytes = int64(2) * 1024 * 1024 * 1024 // 2GiB

This is not a bandwidth throttle. It is a coarse defence-in-depth limit alongside the host allowlist. The allowlist answers where a build may connect. The byte budget catches a different class of mistake: a build that is allowed to talk to a real service, then transfers an implausible amount of data through that permission.

The proxy’s behaviour at the limit was simple:

if not await budget.try_consume_bytes(len(chunk)):
    log.warning("per-run byte budget exceeded - closing connection")
    break

The client does not receive a beautifully typed explanation of that policy. It sees a connection that ended. In this incident, the useful diagnosis existed in the proxy logs while pip produced an incomplete-download error. The policy was useful. The diagnostic boundary was not good enough.

It would have been easy to respond by raising the ceiling. That would have made the build pass and kept an unnecessary GPU stack in an API image. It would also have made the build’s actual network cost less visible. A limit is most useful when it makes an unexpected cost concrete enough to investigate.

torch was not one wheel

The API has a straightforward dependency pin:

torch==2.7.1

There is nothing in that line saying CUDA. That is the problem.

For this platform, ordinary PyPI resolution selected an approximately 821 MB CUDA-enabled Torch wheel and about 2.16 GB of separate nvidia-*-cu12 packages. The larger pieces included nvidia-cudnn-cu12 at 571 MB and nvidia-cublas-cu12 at 393 MB, with roughly ten more CUDA runtime packages behind them.

The dependency graph was doing something sensible for a general Linux Torch installation. It was not doing something Cortex needed.

EmbeddingService chooses the device at runtime:

self.device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

There is no GPU requirement in the Cortex API image build, and no CUDA-specific test path in the API test suite. The one training test that needs real GPU and QLoRA infrastructure is explicitly excluded from this pipeline. Installing the CUDA chain was pure cost here: network bytes, image bytes, extraction time, and a larger place for future vulnerabilities to hide.

The fix installed the matching CPU-only wheel first, from PyTorch’s CPU index. The ordinary requirements install then sees the exact version already installed and resolves the rest of the application dependencies from PyPI as before.

RUN pip install --upgrade pip \
    && pip install torch==2.7.1 --index-url https://download.pytorch.org/whl/cpu \
    && pip install -r requirements.txt

The CPU wheel is approximately 200 MB. It pulls neither the NVIDIA packages nor Triton. That brought the install from about 3.0 GB of package downloads to comfortably below the 2 GiB budget.

flowchart TD
  requirement["torch==2.7.1"] --> pypi["ordinary PyPI resolution"]
  pypi --> torchGpu["CUDA-enabled Torch wheel<br/>about 821 MB"]
  torchGpu --> cudnn["nvidia-cudnn-cu12<br/>571 MB"]
  torchGpu --> cublas["nvidia-cublas-cu12<br/>393 MB"]
  torchGpu --> otherCuda["other nvidia-*-cu12 packages"]
  cudnn --> budget["2 GiB egress budget"]
  cublas --> budget
  otherCuda --> budget
  requirement --> cpuIndex["PyTorch CPU wheel index"]
  cpuIndex --> torchCpu["CPU-only Torch wheel<br/>about 200 MB"]

There are two separate network details inside that small Dockerfile change. download.pytorch.org provides the CPU wheel index, while the actual wheel download redirects to download-r2.pytorch.org. Both belong in the build declaration. Allowlisting the index alone would make package resolution look successful and fail on the larger transfer that follows.

This is the first version of the rule: a package has a network shape. An index endpoint, a redirected file endpoint, the bytes in the distribution, and conditional dependencies are all part of what pip install means in a particular environment.

The image-build node needed to say so

The previous engineering note described why Canopy compiles pipeline TypeScript into a graph before it executes work. Egress is one of the facts carried by that graph. An exec step can declare the hosts it needs. A platform image-build operation can too.

Cortex initially relied on the platform’s automatically injected image-build node. That was enough for an ordinary build with registry access. It was not enough for a Dockerfile that pre-bakes a real embedding model from an external service. The automatic node had no Cortex-specific list of model hosts to enforce.

The pipeline now creates that operation explicitly, after its lint and test inputs:

step.buildImage({
  inputs: [dockerLint, test],
  egress: [
    'huggingface.co',
    '*.aws.cdn.hf.co',
    'cas-server.xethub.hf.co',
    'download.pytorch.org',
    'download-r2.pytorch.org',
    'deb.debian.org',
  ],
});

This is not a claim that a static graph can predict every socket a future version of an arbitrary package manager will open. It cannot. It is a statement of what this known build requests now, where the platform can validate it before dispatch and where a reviewer can inspect it without reconstructing it from a failed container.

The incident was a useful limit of the abstraction. The graph could know that Cortex requested huggingface.co. Running the real model download taught us that this was incomplete. That did not make the declaration less valuable. It gave the declaration better facts.

A hostname was not a capability

Fixing the Torch install exposed the next failure. The Dockerfile installs git, because one API service shells out to it for clone, checkout, diff, and apply work. The Python slim image does not include it.

Adding deb.debian.org to the egress list was the obvious first fix. It did not work.

The base image’s Debian sources used this URL:

http://deb.debian.org/debian

The sandbox proxy is a conventional HTTP CONNECT forward proxy. A build client asks it to open a tunnel, then begins TLS through the tunnel. By default, it permits CONNECT only to port 443. The host allowlist and the port allowlist are deliberately independent checks.

So deb.debian.org being allowed did not mean that deb.debian.org:80 could traverse the sandbox. It could not. apt surfaced that as a temporary DNS-resolution failure, which was another client-level description of a network policy rather than the policy itself.

The Dockerfile now changes Debian’s configured source from HTTP to HTTPS before running apt:

RUN sed -i 's|http://deb.debian.org|https://deb.debian.org|g' /etc/apt/sources.list.d/debian.sources \
    && apt-get update && apt-get install -y --no-install-recommends git \
    && rm -rf /var/lib/apt/lists/*

Opening general port-80 traffic for one package-install step would have treated the proxy’s boundary as an inconvenience. Debian serves the same repository over HTTPS. Using that path preserved the intended policy and made the build’s behaviour more conventional at the same time.

There is a broader distinction worth keeping. A hostname is not a complete network capability. Protocol and port matter. Redirects matter. The amount transferred matters. A policy that records only a hostname is a useful start, but it is not the entire story.

flowchart TD
  declared["Declared host"] --> protocol["protocol"]
  declared --> port["port"]
  declared --> redirects["redirect targets"]
  declared --> volume["bytes transferred"]
  protocol --> result["what the sandbox can actually carry"]
  port --> result
  redirects --> result
  volume --> result

One model was three hosts

The Cortex image pre-bakes its default embedding model, intfloat/multilingual-e5-large, during the Docker build. This is intentional. The Dockerfile stores it in the same Hugging Face cache the service uses at runtime, then sets HF_HUB_OFFLINE=1 after the bake. A normal container start can read the cached default model without a network freshness check.

That is a good runtime boundary. It also means the image build must perform the real download.

At first the build declared huggingface.co and *.aws.cdn.hf.co. That was based on the visible metadata and redirected model-download path. It was not the full path used by the current client.

Tracing a successful download with debug logging found a third host: cas-server.xethub.hf.co. The repository records it as the Xet content-addressed storage backend’s reconstruction API. The exact application-level model download needed all three pieces:

flowchart TD
  build["Cortex Docker build"] --> model["intfloat/multilingual-e5-large"]
  model --> hf["huggingface.co<br/>metadata and redirects"]
  model --> cdn["*.aws.cdn.hf.co<br/>regional CDN/storage traffic"]
  model --> cas["cas-server.xethub.hf.co<br/>Xet CAS reconstruction API"]

That is not an indictment of Hugging Face’s storage design. It is a description of what a client observed while retrieving a large model through modern storage infrastructure. A human asks for one model by one repository name. The client may use metadata, redirects, a regional storage endpoint, and content-addressed reconstruction behind the name.

An allowlisted environment needs to care about that distinction. “Allow Hugging Face” is not a hostname. huggingface.co is not necessarily where every byte of a model comes from. Conversely, naming a few observed hosts should not be presented as a universal model of every Hugging Face download. It is the traced network shape of this model and this build at this point in time.

Construction should not download a model

The image-build failures were not the first sign that Cortex had made network behaviour too implicit.

Earlier that morning, pipeline tests had begun reaching Hugging Face while they were only constructing services. DatasetChunkingService constructed a document-ingestion service. That service obtained an embedding service. EmbeddingService.__init__ immediately loaded a tokenizer and model.

Several paths could therefore cause a model download before any embedding was requested: document ingestion, JSONL ingestion, retrieval, embedding backfill, and incident-memory work all had service factories that eventually obtained the cached embedder. A test that planned to replace one dependency with a stub could have already reached the network before the replacement happened.

The first repair made DatasetChunkingService defer its document-ingestion dependency. The more general repair made the tokenizer and model lazy:

@property
def model(self):
    self._ensure_loaded()
    return self._model

The device decision remains in service construction. The expensive model load happens only when embed_texts() actually needs the tokenizer or model. A fixture that genuinely means to test the real embedder now explicitly accesses .model, rather than confusing construction with reachability.

Lazy loading is not universally correct. Some systems choose an eager startup cost so a health check can prove a dependency is ready. Cortex already has a separate startup-warming control for that decision. What was wrong here was making unrelated object construction silently mean “download a large model from the internet.”

The sandbox made that coupling impossible to ignore. Offline development and tests benefited from the same correction.

The dependency is larger than the import

This was not a story about PyTorch being large. PyTorch was large, but that is too shallow a conclusion.

The useful unit of analysis was not an import line or a package name. It was the dependency’s actual shape:

  • where its metadata comes from
  • where its bytes are served from after redirects
  • which platform-specific dependencies its resolver selects
  • which protocols and ports those services use
  • how much data the path transfers
  • whether it is needed at image build, runtime, or only for a particular operation
  • whether constructing an object begins the network activity

The Cortex build had hidden answers to every one of those questions. Ordinary networking would have let it continue quietly: download several unnecessary gigabytes, carry an oversized image, acquire a list of external services by accident, and make tests depend on a model download because a service object existed.

The sandbox did not cause those properties. It exposed them.

There is still work to do on the platform side. A byte-budget refusal should be visible to the build client as a byte-budget refusal, not merely as an SSL EOF. Host declarations do not currently express every redirect, protocol, port, and expected volume as a complete formal contract. They do make the initial request inspectable, and real execution can reveal where that request was incomplete.

That seems like the right direction. The graph from the previous note is not a promise that CI can know everything before it runs. It is a place to record what it knows, reject what it should not run, and keep the facts learned from reality close to the work that produced them.

pip install torch was not one download. intfloat/multilingual-e5-large was not one hostname. A service constructor was not necessarily a local operation.

Those were all true before the build had a boundary. The boundary finally gave us a reason to see them.