# Optimizing very large neural network that is greater than size of GPU memory

**URL:** <https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427>\
**Category:** Discussion Zone\
**Tags:** ai, neural-network, gpu, tensorflow, qow, researcher\
**Created:** [August 21, 2018, 9:34pm UTC](https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427 "2018-08-21T21:34:02Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![wirawan0](https://ask.cyberinfrastructure.org/user_avatar/ask.cyberinfrastructure.org/wirawan0/32/362_2.png) [@wirawan0](https://ask.cyberinfrastructure.org/u/wirawan0)\
**Post date:** [August 21, 2018, 9:34pm UTC](https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427/1 "2018-08-21T21:34:02Z")

</div>

I encountered a problem with extremely large neural network that was created in KERAS, using Tensorflow backend. The memory footprint in one of the layer is already bigger than the size of current GPU memory (it has just over 4 billion parameters – which, using float32, translates to a matrix 16 GB alone). Is there a neural network implementation that can get around this problem? For example, will TensorFlow accommodate such a case? There are papers/discussions out there on handling matrix multiply that is greater than GPU’s memory size (basically that is done by tiling the matrix and stream the data). Also there is a paper on virtual Deep NN that claims to be transparent to end-user as far as the use of CPU & GPU memory:

“Training Deeper Models by GPU Memory Optimization on TensorFlow”  
[http://learningsys.org/nips17/assets/papers/paper\_18.pdf](http://learningsys.org/nips17/assets/papers/paper_18.pdf)

" vDNN: Virtualized Deep Neural Networks for Scalable, Memory-Efficient Neural Network Design"

> **[1602.08124.pdf](https://arxiv.org/pdf/1602.08124.pdf)**
>
> 2.52 MB

“How to Train a Very Large and Deep Model on One GPU?”

> **[How to Train a Very Large and Deep Model on One GPU?](https://medium.com/syncedreview/how-to-train-a-very-large-and-deep-model-on-one-gpu-7b7edfe2d072)**
>
> Problem: GPU memory limitation

but my question is simply: is this doable using current neural network implementation? Tensorflow claims to support parallel computation, multiple GPU, etc. Will Tensorflow accommodate cases like that one above without choking?

---

<div class="post-metadata">

**Author:** ![raminder](https://ask.cyberinfrastructure.org/letter_avatar_proxy/v4/letter/r/58956e/32.png) [@raminder](https://ask.cyberinfrastructure.org/u/raminder)\
**Post date:** [August 28, 2018, 1:39am UTC](https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427/2 "2018-08-28T01:39:53Z")

</div>

Thanks for sharing the paper and reading. It’s really interesting. While doing research on this, I found [https://medium.com/tensorflow/fitting-larger-networks-into-memory-583e3c758ff9](https://medium.com/tensorflow/fitting-larger-networks-into-memory-583e3c758ff9) which may be useful for you.

---

<div class="post-metadata">

**Author:** ![cec5550](https://ask.cyberinfrastructure.org/letter_avatar_proxy/v4/letter/c/e8c25b/32.png) [@cec5550](https://ask.cyberinfrastructure.org/u/cec5550)\
**Post date:** [June 13, 2023, 1:03pm UTC](https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427/3 "2023-06-13T13:03:28Z")

</div>

I would recommend to try the unified memory option in tensorflow by setting the env var `TF_FORCE_UNIFIED_MEMORY` to `true` .  
It allows the memory overflow to fall back on the CPU or alternatively to try half precision which should lower your memory footprint.

---

<div class="post-metadata">

**Author:** ![jfossot](https://ask.cyberinfrastructure.org/user_avatar/ask.cyberinfrastructure.org/jfossot/32/932_2.png) [@jfossot](https://ask.cyberinfrastructure.org/u/jfossot)\
**Post date:** [August 17, 2023, 3:10pm UTC](https://ask.cyberinfrastructure.org/t/optimizing-very-large-neural-network-that-is-greater-than-size-of-gpu-memory/427/4 "2023-08-17T15:10:30Z")

</div>

OR use a system set up for ML/AI Big data. Go through the ACCESS Resource List. You will find many that meet your needs  
[https://allocations.access-ci.org/resources](https://allocations.access-ci.org/resources)
