Education and research
How universities and research institutes use NVIDIA AI: shared GPU clusters and national supercomputers, fair scheduling across labs, campus language assistants, robot and science simulation, and teaching resources, with research data governance treated as a design issue.
The problem
Research groups across a university compete for the same scarce GPUs. Without a shared service, labs buy their own servers that sit idle between grant deadlines, while other groups wait weeks for cloud budget or queue time on a central cluster that was sized for CPU jobs.
Data adds its own constraints. Clinical, genomic and survey datasets arrive with consent terms, export rules and data use agreements, and industry partners may supply data that must stay inside a secure enclave. Results must be reproducible years later, which is hard when software stacks and container images change every few months.
Teaching has a separate gap: courses in many disciplines, not only computer science, need hands-on GPU material, and smaller institutions and even whole countries may have no accelerated computing of their own.1
The approach
Many campuses now run one shared AI platform instead of many lab servers. The University of Pennsylvania built a shared platform called Betty on DGX SuperPOD, Stanford runs its Marlowe DGX SuperPOD alongside the existing Sherlock cluster for all seven of its schools, and Cal Poly opened an AI factory of four DGX B200 systems for students, faculty and regional partners. At national level, Denmark's Gefion supercomputer, a DGX SuperPOD run by DCAI, opened with a pilot phase for selected university, startup and public research projects.
Sharing needs a scheduler that enforces fairness. EPFL's research computing platform uses Run:ai on Kubernetes to give each lab or project a guaranteed GPU quota and to lend unused GPUs to others. On top of the cluster, institutions build services: the University of Florida runs its NaviGator chatbots with NIM microservices on the HiPerGator supercomputer, and University of Maryland researchers use Isaac to build virtual homes where robots practice household tasks. For teaching, NVIDIA offers DLI Teaching Kits for coursework, and its Academic Grant Program has funded projects such as the Maryland robotics work.
Smaller departments often do better with an existing Slurm cluster, a national research computing allocation or cloud credits than with buying DGX systems, and many statistics and simulation codes still run well on CPUs. A dedicated GPU platform makes sense when demand is steady across many groups and someone is funded to operate it.123
Conceptual architecture
Diagram as a list
Applications & solutions
- Researchers, students and partner projectsSubmit notebooks, training runs, simulations and courseworkConnects to Allocation and accounts (quotas per lab, course or grant)
- Simulation and robot learning (Isaac)Trains and tests robot behavior in virtual environments before lab trialsConnects to Fair-share GPU scheduler (Run:ai on Kubernetes, or Slurm)
Inference & runtime software
- Campus AI services (assistants served with NIM)Gives students and staff model-backed tools without each lab hosting its ownConnects to Researchers, students and partner projects
Operations & orchestration
- Allocation and accounts (quotas per lab, course or grant)Turns the institution's access policy into quotas and prioritiesConnects to Fair-share GPU scheduler (Run:ai on Kubernetes, or Slurm)
- Fair-share GPU scheduler (Run:ai on Kubernetes, or Slurm)Guarantees each group its share and lends idle GPUs to othersConnects to Shared campus cluster or national supercomputer (DGX SuperPOD)
- Governed research data store and secure enclaveHolds sensitive datasets under their use agreements and logs accessConnects to Shared campus cluster or national supercomputer (DGX SuperPOD)
Accelerated computing
- Shared campus cluster or national supercomputer (DGX SuperPOD)Runs training, inference and simulation jobs
Programs & resources
- Teaching kits, courses and academic grantsSupplies course material and hardware or grants for selected researchConnects to Researchers, students and partner projects
Technologies and their roles
NVIDIA DGX12
Shared campus and national clusters
Penn, Stanford and Cal Poly run DGX systems as shared research platforms, and Denmark's Gefion is a DGX SuperPOD.
NVIDIA Run:ai34
Fair GPU sharing across labs
EPFL uses Run:ai on Kubernetes to give each lab a guaranteed GPU quota and lend idle capacity to other groups.
NVIDIA NIM2
Campus language services
The University of Florida serves its NaviGator chatbots with NIM microservices on HiPerGator.
NVIDIA Isaac2
Robotics research in simulation
University of Maryland researchers use Isaac to create virtual homes for training household robots.
What you need first
- An access policy that says who gets compute: an allocation committee, quotas per lab or course, and priority rules for deadlines
- Research data management plans that cover consent, export controls and data use agreements
- A research computing team with Kubernetes or Slurm, storage and GPU driver skills
- Integration with campus identity so access follows projects and grants
- Data center power and cooling for dense GPU systems, or a hosting partner
- Versioned container environments so published results can be reproduced
- A budget for operations and hardware refresh, not only the initial purchase
Risks and how to reduce them
- One group monopolizing a shared cluster3
- Use guaranteed quotas with fair-share borrowing and preemption, publish the allocation policy, and review usage by group each term.
- Sensitive research data exposed on a shared system
- Run projects with sensitive data in isolated namespaces or enclaves, enforce data use agreements technically, and log every access.
- Results that cannot be reproduced
- Pin container images, model versions and random seeds, and archive them with each publication.
- Expensive hardware left underused
- Measure demand with cloud credits or national allocations before buying, and open spare capacity to partner institutions.
- Campus assistants giving wrong answers on policy, grades or wellbeing
- Ground assistants in official documents, show sources, and route sensitive questions to staff.
Documented examples
Danish Centre for AI Innovation (DCAI) · Research and AI infrastructure (sovereign AI)
DCAI Gefion: Denmark's sovereign AI supercomputer on DGX SuperPOD
Gefion, operated by the Danish Centre for AI Innovation, is an NVIDIA DGX SuperPOD with 1,528 H100 GPUs, funded by the Novo Nordisk Foundation and Denmark's export and investment fund. It went live in October 2024, placed 21st on the November 2024 TOP500 list, and now serves researchers and companies.
In production
Related
Sources
Thank you. Your correction was sent.
The editors check it against the sources. If you left an email address, they may reply about it.