GPU Rentals

Decentralized GPU sharing platform for distributed model training.

Ray
Docker
Next.js
TypeScript
WebSocket
Distributed Computing

Problem

GPUs sit idle in some machines while other developers can't access compute for training — there is no simple way to share spare capacity across a network of nodes.

Why I built it

To let users rent out spare GPUs and run distributed training tasks across them, with allocation based on node capacity and results aggregated after execution.

Architecture

  1. Node registration

    GPU owners register their nodes with advertised capacity so the scheduler knows what compute is available.

  2. Capacity-based allocation

    Incoming training tasks are allocated across nodes based on capacity, distributing work instead of overloading a single machine.

  3. Distributed execution

    Ray coordinates distributed execution across nodes, containerized with Docker for reproducible environments.

  4. Aggregation

    Per-node results are aggregated post-execution and surfaced through a Next.js frontend with WebSocket updates.

Engineering challenges

  • Allocating tasks fairly across heterogeneous nodes with very different capacity.
  • Keeping distributed runs observable end to end, from scheduling to aggregation.

Lessons learned

  • Scheduling is the heart of a sharing platform — the execution engine matters less than fair, transparent allocation.
  • Containerized workers make heterogeneous hardware behave predictably.