# \#gpu

**URL:** https://ask.cyberinfrastructure.org/tag/gpu/51.md

[Latest](https://ask.cyberinfrastructure.org/latest.md) · [Categories](https://ask.cyberinfrastructure.org/categories.md) · [Tags](https://ask.cyberinfrastructure.org/tags.md)

---

## [Multi-user ArcGIS Pro for sensitive research data—what's working at scale?](https://ask.cyberinfrastructure.org/t/multi-user-arcgis-pro-for-sensitive-research-data-whats-working-at-scale/4380)

<div class="topic-metadata">

**Author:** [@mitchellxh](https://ask.cyberinfrastructure.org/u/mitchellxh)\
**Replies:** 1\
**Last updated:** [January 5, 2026, 11:42pm UTC](https://ask.cyberinfrastructure.org/t/multi-user-arcgis-pro-for-sensitive-research-data-whats-working-at-scale/4380 "2026-01-05T23:42:12Z")

</div>

We’re trying to provide shared ArcGIS Pro access to researchers working with high-risk data (PHI-adjacent, contractually confidential). Currently running persistent Azure VMs (bundles Windows OS license unlike AWS). Woul…

---

## [How to run a research deepfake model on Jetstream2](https://ask.cyberinfrastructure.org/t/how-to-run-a-research-deepfake-model-on-jetstream2/4035)

<div class="topic-metadata">

**Author:** [@juanjo.garciamesa](https://ask.cyberinfrastructure.org/u/juanjo.garciamesa)\
**Replies:** 2\
**Last updated:** [March 10, 2025, 9:03pm UTC](https://ask.cyberinfrastructure.org/t/how-to-run-a-research-deepfake-model-on-jetstream2/4035 "2025-03-10T21:03:47Z")

</div>

I recently was able to deploy an instance in Jetstream2 and run a deepfake project that focuses on individual recognition in paper wasps. I composed a document with the steps I followed and I thought it would be of inter…

---

## [Multinode distributed gpu](https://ask.cyberinfrastructure.org/t/multinode-distributed-gpu/3072)

<div class="topic-metadata">

**Author:** [@sladenheim](https://ask.cyberinfrastructure.org/u/sladenheim)\
**Replies:** 9\
**Last updated:** [May 30, 2024, 5:46pm UTC](https://ask.cyberinfrastructure.org/t/multinode-distributed-gpu/3072 "2024-05-30T17:46:16Z")

</div>

I am interested in knowing which of the ACCESS systems allow for multinode distributed GPU applications (e.g., I have 2 compute nodes each with 4 gpus for a total of 8 gpus). I am aware that Delta has this capability but…

---

## [Resources for GPU and Google JAX programming](https://ask.cyberinfrastructure.org/t/resources-for-gpu-and-google-jax-programming/3133)

<div class="topic-metadata">

**Author:** [@sam\_dt](https://ask.cyberinfrastructure.org/u/sam_dt)\
**Replies:** 0\
**Last updated:** [May 17, 2024, 2:27pm UTC](https://ask.cyberinfrastructure.org/t/resources-for-gpu-and-google-jax-programming/3133 "2024-05-17T14:27:23Z")

</div>

When I first started my project in Google JAX, I knew that that Google JAX was a pretty new coding library made for GPUs and that examples and teaching materials would be scarce for that. And in addition, I was pretty ne…

---

## [Prevent CPU-only Jobs running on GPU nodes](https://ask.cyberinfrastructure.org/t/prevent-cpu-only-jobs-running-on-gpu-nodes/2696)

<div class="topic-metadata">

**Author:** [@DelilahM](https://ask.cyberinfrastructure.org/u/DelilahM)\
**Replies:** 6\
**Last updated:** [April 26, 2024, 7:59pm UTC](https://ask.cyberinfrastructure.org/t/prevent-cpu-only-jobs-running-on-gpu-nodes/2696 "2024-04-26T19:59:26Z")

</div>

Hi All, Does anyone have any experience configuring SLURM that prevents cpu-only jobs running on GPU nodes? TIA Delilah

---

## [Common Issue RuntimeError: CUDA error](https://ask.cyberinfrastructure.org/t/common-issue-runtimeerror-cuda-error/2943)

<div class="topic-metadata">

**Author:** [@ShrutiDongare](https://ask.cyberinfrastructure.org/u/ShrutiDongare)\
**Replies:** 1\
**Last updated:** [November 2, 2023, 8:20pm UTC](https://ask.cyberinfrastructure.org/t/common-issue-runtimeerror-cuda-error/2943 "2023-11-02T20:20:46Z")

</div>

I am running into ‘RuntimeError: CUDA error: no kernel image is available for execution on the device’ multiple times. I tried to manually configure Pytorch and its compatible cuda version. The versions that did not work…

---

## [User-Initiated GPU Monitoring during 'Draining' State](https://ask.cyberinfrastructure.org/t/user-initiated-gpu-monitoring-during-draining-state/2965)

<div class="topic-metadata">

**Author:** [@ShrutiDongare](https://ask.cyberinfrastructure.org/u/ShrutiDongare)\
**Replies:** 2\
**Last updated:** [October 2, 2023, 1:05pm UTC](https://ask.cyberinfrastructure.org/t/user-initiated-gpu-monitoring-during-draining-state/2965 "2023-10-02T13:05:35Z")

</div>

I frequently encounter a situation where, after submitting a job to the HPC cluster, the job state remains queued for an extended period. Upon reaching out to the system administrator, I was informed that GPUs were in a …

---

## [Inter-node and intra-node GPU communication](https://ask.cyberinfrastructure.org/t/inter-node-and-intra-node-gpu-communication/2954)

<div class="topic-metadata">

**Author:** [@ShrutiDongare](https://ask.cyberinfrastructure.org/u/ShrutiDongare)\
**Replies:** 2\
**Last updated:** [October 1, 2023, 3:51am UTC](https://ask.cyberinfrastructure.org/t/inter-node-and-intra-node-gpu-communication/2954 "2023-10-01T03:51:02Z")

</div>

What tools or methods would you recommend for measuring both inter-node GPU communication and intra-node GPU communication in a high-performance computing or distributed computing environment?

---

## [Multi-instance GPU](https://ask.cyberinfrastructure.org/t/multi-instance-gpu/1710)

<div class="topic-metadata">

**Author:** [@wirawan0](https://ask.cyberinfrastructure.org/u/wirawan0)\
**Replies:** 2\
**Last updated:** [October 1, 2023, 2:05am UTC](https://ask.cyberinfrastructure.org/t/multi-instance-gpu/1710 "2023-10-01T02:05:53Z")

</div>

Hi folks, the new Ampere A100 GPU supports “multi-instance GPU”, i.e. allowing a GPU to be partitioned so multiple people can use the GPU to maximize throughout of computations. Other than Ampere A100, any other cards or…

---

## [Examine specific lines of code for memory profiling](https://ask.cyberinfrastructure.org/t/examine-specific-lines-of-code-for-memory-profiling/2963)

<div class="topic-metadata">

**Author:** [@shrek](https://ask.cyberinfrastructure.org/u/shrek)\
**Replies:** 1\
**Last updated:** [September 30, 2023, 7:50pm UTC](https://ask.cyberinfrastructure.org/t/examine-specific-lines-of-code-for-memory-profiling/2963 "2023-09-30T19:50:44Z")

</div>

How can I obtain detailed insights for analyzing hardware memory usage in fine detail within a simple Python function as shown below, especially when examining specific lines of code contributing to memory growth, simila…

---

## [Page-level memory management for NVIDIA GPUs](https://ask.cyberinfrastructure.org/t/page-level-memory-management-for-nvidia-gpus/2947)

<div class="topic-metadata">

**Author:** [@ShrutiDongare](https://ask.cyberinfrastructure.org/u/ShrutiDongare)\
**Replies:** 0\
**Last updated:** [September 21, 2023, 9:39pm UTC](https://ask.cyberinfrastructure.org/t/page-level-memory-management-for-nvidia-gpus/2947 "2023-09-21T21:39:43Z")

</div>

My objective is to monitor memory access patterns for generative AI workloads and subsequently develop optimization strategies for managing memory costs across various memory tiers. One example of memory tiers would be G…

---

## [Optimizing very large neural network that is greater than size of GPU memory](https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427)

<div class="topic-metadata">

**Author:** [@wirawan0](https://ask.cyberinfrastructure.org/u/wirawan0)\
**Replies:** 3\
**Last updated:** [August 17, 2023, 3:10pm UTC](https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427 "2023-08-17T15:10:30Z")

</div>

I encountered a problem with extremely large neural network that was created in KERAS, using Tensorflow backend. The memory footprint in one of the layer is already bigger than the size of current GPU memory (it has just…

---

## [Why is GPU not found after installing PyTorch via Anaconda?](https://ask.cyberinfrastructure.org/t/why-is-gpu-not-found-after-installing-pytorch-via-anaconda/2621)

<div class="topic-metadata">

**Author:** [@wwarr](https://ask.cyberinfrastructure.org/u/wwarr)\
**Replies:** 4\
**Last updated:** [August 17, 2023, 3:03pm UTC](https://ask.cyberinfrastructure.org/t/why-is-gpu-not-found-after-installing-pytorch-via-anaconda/2621 "2023-08-17T15:03:44Z")

</div>

I’ve installed PyTorch in an Anaconda environment but can’t find the GPU. What is going wrong, why, and how can I fix it?

---

## [How to use multiple GPU node to train/fine-tune large language models?](https://ask.cyberinfrastructure.org/t/how-to-use-multiple-gpu-node-to-train-fine-tune-large-language-models/2816)

<div class="topic-metadata">

**Author:** [@sohil.shrestha](https://ask.cyberinfrastructure.org/u/sohil.shrestha)\
**Replies:** 3\
**Last updated:** [May 31, 2023, 7:00pm UTC](https://ask.cyberinfrastructure.org/t/how-to-use-multiple-gpu-node-to-train-fine-tune-large-language-models/2816 "2023-05-31T19:00:57Z")

</div>

I am using UT’s Texas Advanced Computing Center (TACC)'s resources to fine tune large language models. However, I am able to fine-tune it with 1-batch size making the fine-tuning extremely slow and unstable. My understa…

---

## [Recommendations from HPC/ML specialists and enthusiasts](https://ask.cyberinfrastructure.org/t/recommendations-from-hpc-ml-specialists-and-enthusiasts/2757)

<div class="topic-metadata">

**Author:** [@shakhizat](https://ask.cyberinfrastructure.org/u/shakhizat)\
**Replies:** 1\
**Last updated:** [April 22, 2023, 1:25pm UTC](https://ask.cyberinfrastructure.org/t/recommendations-from-hpc-ml-specialists-and-enthusiasts/2757 "2023-04-22T13:25:36Z")

</div>

Dear Ask Cyberinfrastructure community, As a small non-profit academic organization with Nvidia DGX servers, we are seeking recommendations from HPC/AI specialists and enthusiasts worldwide. We would like to know how yo…

---

## [Rutgers University is hiring a Director of Research Support!](https://ask.cyberinfrastructure.org/t/rutgers-university-is-hiring-a-director-of-research-support/2713)

<div class="topic-metadata">

**Author:** [@JanaeBaker](https://ask.cyberinfrastructure.org/u/JanaeBaker)\
**Replies:** 0\
**Last updated:** [March 20, 2023, 10:06pm UTC](https://ask.cyberinfrastructure.org/t/rutgers-university-is-hiring-a-director-of-research-support/2713 "2023-03-20T22:06:36Z")

</div>

The Rutgers’ Office of Advanced Research Computing (OARC) is on the look for a Director of Research Support to join our growing team! Rutgers is a leading national research university and the State of New Jersey’s preem…

---

## [Systems Administrators in Pittsburgh: Servers, Clusters and Supercomputers for Computational Biochemistry](https://ask.cyberinfrastructure.org/t/systems-administrators-in-pittsburgh-servers-clusters-and-supercomputers-for-computational-biochemistry/2605)

<div class="topic-metadata">

**Author:** [@DESRESCareers](https://ask.cyberinfrastructure.org/u/DESRESCareers)\
**Replies:** 0\
**Last updated:** [November 7, 2022, 6:12am UTC](https://ask.cyberinfrastructure.org/t/systems-administrators-in-pittsburgh-servers-clusters-and-supercomputers-for-computational-biochemistry/2605 "2022-11-07T06:12:15Z")

</div>

Exceptional sysadmins sought to manage systems infrastructure, networks, and data center operations at a state-of-the-art facility in Pittsburgh, PA. Successful hires will be responsible for bring-up and ongoing mainten…

---

## [Systems Administrators in New York City: Servers, Clusters and Supercomputers for Computational Biochemistry](https://ask.cyberinfrastructure.org/t/systems-administrators-in-new-york-city-servers-clusters-and-supercomputers-for-computational-biochemistry/2350)

<div class="topic-metadata">

**Author:** [@DESRESCareers](https://ask.cyberinfrastructure.org/u/DESRESCareers)\
**Replies:** 0\
**Last updated:** [May 18, 2022, 11:42am UTC](https://ask.cyberinfrastructure.org/t/systems-administrators-in-new-york-city-servers-clusters-and-supercomputers-for-computational-biochemistry/2350 "2022-05-18T11:42:31Z")

</div>

Exceptional sysadmins sought to manage systems infrastructure, networks, and information security for a New York–based interdisciplinary research group. Successful hires will be responsible for our servers and cluster e…

---

## [Why can't TensorFlow find a GPU on Cheaha?](https://ask.cyberinfrastructure.org/t/why-cant-tensorflow-find-a-gpu-on-cheaha/2507)

<div class="topic-metadata">

**Author:** [@wwarr](https://ask.cyberinfrastructure.org/u/wwarr)\
**Replies:** 1\
**Last updated:** [September 16, 2022, 4:36pm UTC](https://ask.cyberinfrastructure.org/t/why-cant-tensorflow-find-a-gpu-on-cheaha/2507 "2022-09-16T16:36:11Z")

</div>

I’m using TensorFlow on Cheaha, and my code isn’t using the GPU or I can’t locate the GPU or I’m receiving an error about missing GPUs. Why is this occurring and what can I do about it?

---

## [Sys Amin Job Opening at Rutgers University](https://ask.cyberinfrastructure.org/t/sys-amin-job-opening-at-rutgers-university/2389)

<div class="topic-metadata">

**Author:** [@JanaeBaker](https://ask.cyberinfrastructure.org/u/JanaeBaker)\
**Replies:** 0\
**Last updated:** [June 10, 2022, 6:18pm UTC](https://ask.cyberinfrastructure.org/t/sys-amin-job-opening-at-rutgers-university/2389 "2022-06-10T18:18:28Z")

</div>

Dear Colleagues, The Rutgers University Office of Advanced Research Computing (OARC) is seeking a Systems Administrator to join our team. We are a large and diverse team working to create an outstanding environment for…

---

## [Digital Research Consultant at University of California Los Angeles (UCLA)](https://ask.cyberinfrastructure.org/t/digital-research-consultant-at-university-of-california-los-angeles-ucla/2373)

<div class="topic-metadata">

**Author:** [@Jmartinez](https://ask.cyberinfrastructure.org/u/Jmartinez)\
**Replies:** 0\
**Last updated:** [June 2, 2022, 10:42pm UTC](https://ask.cyberinfrastructure.org/t/digital-research-consultant-at-university-of-california-los-angeles-ucla/2373 "2022-06-02T22:42:40Z")

</div>

We’re pleased to share a new opportunity at UCLA with the Office of Advanced Research Computing organization for a Digital Research Consultant professional within the GIS, Visualization, XR & Modeling group. Join a team…

---

## [New ideas for GPU hardware](https://ask.cyberinfrastructure.org/t/new-ideas-for-gpu-hardware/2317)

<div class="topic-metadata">

**Author:** [@Jack\_Swope](https://ask.cyberinfrastructure.org/u/Jack_Swope)\
**Replies:** 2\
**Last updated:** [April 25, 2022, 6:02pm UTC](https://ask.cyberinfrastructure.org/t/new-ideas-for-gpu-hardware/2317 "2022-04-25T18:02:31Z")

</div>

We are currently looking at new hardware packages to get more GPUs into our systems. We have been buying 1U and 2U rackmount servers with one to multiple GPUs installed. We are possibly looking for blade/chassis soluti…

---

## [Software Engineer for the HTCondor Software Suite](https://ask.cyberinfrastructure.org/t/software-engineer-for-the-htcondor-software-suite/2257)

<div class="topic-metadata">

**Author:** [@LAUREN\_MICHAEL](https://ask.cyberinfrastructure.org/u/LAUREN_MICHAEL)\
**Replies:** 0\
**Last updated:** [March 17, 2022, 7:04pm UTC](https://ask.cyberinfrastructure.org/t/software-engineer-for-the-htcondor-software-suite/2257 "2022-03-17T19:04:48Z")

</div>

Join CHTC’s development team for the HTCondor Software Suite, or help spread the word to those who may be interested! https://jobs.hr.wisc.edu/en-us/job/512630/software-engineer CHTC is hiring a new member of our softw…

---

## [Monitoring GPU usage](https://ask.cyberinfrastructure.org/t/monitoring-gpu-usage/932)

<div class="topic-metadata">

**Author:** [@jma](https://ask.cyberinfrastructure.org/u/jma)\
**Replies:** 4\
**Last updated:** [December 6, 2021, 8:37pm UTC](https://ask.cyberinfrastructure.org/t/monitoring-gpu-usage/932 "2021-12-06T20:37:28Z")

</div>

I’m running a calculation that parallelizes well using both MPI and OpenMP. While reading online, it occurred to me that I might gain speedup with a GPU node. How can I estimate the potential benefit? And is it possib…

---

## [Developing code which can run on AMD and NVIDIA GPUs](https://ask.cyberinfrastructure.org/t/developing-code-which-can-run-on-amd-and-nvidia-gpus/2060)

<div class="topic-metadata">

**Author:** [@david.matthews.1](https://ask.cyberinfrastructure.org/u/david.matthews.1)\
**Replies:** 1\
**Last updated:** [September 23, 2021, 3:03pm UTC](https://ask.cyberinfrastructure.org/t/developing-code-which-can-run-on-amd-and-nvidia-gpus/2060 "2021-09-23T15:03:33Z")

</div>

The clusters I have access to have a mix of AMD and NVIDIA GPUs. How can I develop code which is easy to run with any accelerator?

---

## [Slurm, GPU, CGroups, ConstrainDevices](https://ask.cyberinfrastructure.org/t/slurm-gpu-cgroups-constraindevices/1745)

<div class="topic-metadata">

**Author:** [@stryder](https://ask.cyberinfrastructure.org/u/stryder)\
**Replies:** 2\
**Last updated:** [May 11, 2021, 9:47pm UTC](https://ask.cyberinfrastructure.org/t/slurm-gpu-cgroups-constraindevices/1745 "2021-05-11T21:47:29Z")

</div>

Has any site managed to make slurm + constraindevices + cgroups work well enough that it actually prevents people from manually setting CUDA\_VISIBLE\_DEVICES and double booking processes on GPU? I can see constraindevice…

---

## [Customize modulefile based on the node properties](https://ask.cyberinfrastructure.org/t/customize-modulefile-based-on-the-node-properties/552)

<div class="topic-metadata">

**Author:** [@ktrn](https://ask.cyberinfrastructure.org/u/ktrn)\
**Replies:** 3\
**Last updated:** [May 11, 2021, 9:01pm UTC](https://ask.cyberinfrastructure.org/t/customize-modulefile-based-on-the-node-properties/552 "2021-05-11T21:01:43Z")

</div>

What is the best way to customize a modulefile (used to specify a particular version of a software) based on some properties of the node? For example, if the node has a GPU, I would like to set the PATH environment varia…

---

## [Slurm: Gres vs --gpus configuration, syntax preference on A100](https://ask.cyberinfrastructure.org/t/slurm-gres-vs-gpus-configuration-syntax-preference-on-a100/1652)

<div class="topic-metadata">

**Author:** [@stryder](https://ask.cyberinfrastructure.org/u/stryder)\
**Replies:** 1\
**Last updated:** [November 30, 2020, 7:42pm UTC](https://ask.cyberinfrastructure.org/t/slurm-gres-vs-gpus-configuration-syntax-preference-on-a100/1652 "2020-11-30T19:42:22Z")

</div>

Previously many sites used the GRES syntax for user facing commands to request GPU using slurm. Now with the newer --gpus syntax is there a preference? Can the two co-exist at the same time for users bringing older codes…

---

## [Installing cupy](https://ask.cyberinfrastructure.org/t/installing-cupy/1482)

<div class="topic-metadata">

**Author:** [@antoinearnoud](https://ask.cyberinfrastructure.org/u/antoinearnoud)\
**Replies:** 1\
**Last updated:** [August 26, 2020, 1:44pm UTC](https://ask.cyberinfrastructure.org/t/installing-cupy/1482 "2020-08-26T13:44:56Z")

</div>

Hi, I am using the python module OPT. I would like to use the GPU implementation of some functions. It seems it requires the module cupy. I tried to install cupy (pip install cupy) but I get an error: ERROR: CUDA coul…

---

## [When using a GPU to accelerate NAMD, what are the drawbacks of using a single-precision GPU?](https://ask.cyberinfrastructure.org/t/when-using-a-gpu-to-accelerate-namd-what-are-the-drawbacks-of-using-a-single-precision-gpu/1429)

<div class="topic-metadata">

**Author:** [@jgoodhue](https://ask.cyberinfrastructure.org/u/jgoodhue)\
**Replies:** 2\
**Last updated:** [July 27, 2020, 7:57pm UTC](https://ask.cyberinfrastructure.org/t/when-using-a-gpu-to-accelerate-namd-what-are-the-drawbacks-of-using-a-single-precision-gpu/1429 "2020-07-27T19:57:27Z")

</div>

I see that NAMD has versions for both double precision and single precision GPUs. Single-precision GPUs are attractive because they are less expensive. Can someone say how the lower precision would affect the quality o…

[Next page](https://ask.cyberinfrastructure.org/tag/gpu/51.md?match_all_tags=true&page=1&tags%5B%5D=gpu)
