Google Compute Engine’s e2-micro instances are a staple of the GCP Free Tier, offering 2 vCPUs and roughly 1 GB of RAM (969 MiB usable). They are ideal for lightweight bastions, cron runners, and small test beds. However, if you run memory-sensitive workloads—such as Go builds, package extractions, or continuous process forks—you may have experienced intermittent, baffling SSH connection timeouts and unresponsive terminals.

The root cause is rarely your application alone. By default, GCE provisions a suite of Google guest daemons that consume a disproportionate share of resident memory on small VMs, creating an environment primed for SSH exhaustion.


The Anatomy of the 15% Memory Tax

Running a bare Debian 12 Bookworm instance on an e2-micro without any user workload active reveals an astonishing baseline:

Process / Daemon Command Name Resident RAM (RSS) % of 1 GB RAM
Google OS Config Agent google_osconfig ~61.8 MB ~6.2%
Guest Agent Core Plugin core_plugin ~27.7 MB ~2.8%
Guest Agent Manager google_guest_ag ~27.2 MB ~2.7%
Guest Compat Manager google_guest_co ~19.7 MB ~2.0%
Guest Telemetry Extension guest_telemetry ~18.7 MB ~1.9%
Total Google Guest Footprint ~155.1 MB ~15.6%

On a 16 GB node, a 155 MB footprint is negligible noise (< 1%). On a 1 GB instance without swap space, 15.6% of your physical memory is permanently claimed before your application even boots.


Why OS Config Hoards 62 MB of RAM

Static analysis of the GoogleCloudPlatform/osconfig repository reveals why google_osconfig claims over 60 MB alone:

  1. Embedded Vulnerability Scanners: Rather than calling out to an external CLI tool and exiting, osconfig embeds Google’s osv-scalibr scanner directly in-process. It parses package databases (dpkg, rpm, pip, gem) into memory trees.
  2. Heavy Protobuf & gRPC Slices: Every inventory cycle instantiates fresh gRPC client transports and serializes thousands of package messages into contiguous memory chunks.
  3. Go Runtime Retention: The Go garbage collector runs, but memory marked as free (madvise(MADV_DONTNEED)) is not eagerly returned to the kernel RSS table. It holds its high-water mark indefinitely.

The SSH Starvation Cascade: How OS Login Fails Under Load

When your application executes builds or unpacks files, kernel VFS slab caches (dentry, ext4_inode_cache) grow dynamically. When total physical RAM nears exhaustion, Linux enters severe memory pressure (surfacing in /proc/pressure/memory).

If OS Login (enable-oslogin=TRUE) is active, incoming SSH connections must complete the following chain:

  1. sshd receives a TCP connection on port 22.
  2. sshd executes AuthorizedKeysCommand (/usr/bin/google_authorized_keys).
  3. The PAM stack loads /lib/security/pam_oslogin_login.so.
  4. PAM performs IPC calls or local HTTP round-trips to the GCE metadata server (169.254.169.254) to verify IAM permissions and POSIX mappings.

Under memory stall conditions (PSI some avg10 > 2.0), fork() calls stall, metadata socket round-trips lag, and the SSH client hits its default ConnectTimeout (5–10s). The user sees:

ssh: connect to host 34.xx.xx.xx port 22: Connection timed out
# or
Connection reset by peer / Permission denied (publickey)

Conversely, with standard SSH authorized keys, sshd opens local /home/<user>/.ssh/authorized_keys directly with a single disk read, bypassing external IPC entirely.


Reclaiming Your RAM: Disabling OS Config & OS Login

If you don’t use GCP VM Manager (automated patch management dashboards in the Cloud Console) or IAM-based OS Login, disabling them recovers immediate memory and makes SSH bulletproof.

You can enforce this across all instances in a project with two gcloud commands:

# 1. Disable OS Config and OS Login project-wide
gcloud compute project-info add-metadata \
    --project="my-project" \
    --metadata=enable-osconfig=FALSE,enable-oslogin=FALSE

# 2. Add your static SSH public key to project metadata
gcloud compute project-info add-metadata \
    --project="my-project" \
    --metadata="ssh-keys=tonymet:$(cat ~/.ssh/id_rsa.pub)"

Note: Disabling OS Login allows the standard google-guest-agent to automatically read the project-wide ssh-keys metadata and write /home/tonymet/.ssh/authorized_keys on every instance.

Option 2: Inside the Operating System

If you prefer to disable the services directly inside the VM:

# Stop and mask the OS Config service
sudo systemctl stop google-osconfig-agent
sudo systemctl disable --now google-osconfig-agent
sudo systemctl mask google-osconfig-agent

# Optional: Completely remove the package
sudo apt-get purge -y google-osconfig-agent

Telemetry Comparison: Before vs. After

Using our synthetic stress harness on an e2-micro running Debian 12 (pinned 250 MB heap anchor, 10,000 file/dentry cycles, and continuous process execution churn):

Metric Agents Active (Control) Agents Stopped (Treatment) Improvement
Baseline Resident RAM ~325 MB used ~170 MB used ~155 MB recovered
Max Memory Stall (some avg10) 2.37 0.62 -73.8% pressure stall reduction
Max Full Memory Stall (full avg10) 1.74 0.38 -78.2% full stall reduction
SSH Probe Failure Rate 0 timeouts 0 timeouts Sub-second latency

Key Takeaways

  1. Know Your Guest Footprint: Google guest agents provide cloud-native conveniences, but their ~155 MB overhead is disproportionately taxing on free-tier 1 GB instances.
  2. OS Config is Optional: Disabling it only removes centralized fleet patching and inventory dashboards from the Cloud Console; basic VM operations, networking, and standard updates remain 100% intact.
  3. Simpler Auth is Resilient Auth: On constrained systems, local authorized keys eliminate IPC round-trips to the metadata server, ensuring reliable console access even when the system is under stress.