Context: Let's look at a simple example of why vanishing and exploding gradients occur in RNNs. Consider a univariate version of RNN with the following update rules ๐‘ง ( ๐‘ก ) = ๐‘ข ๐‘ฅ ( ๐‘ก ) + ๐‘ค โ„Ž ( ๐‘ก โˆ’ 1 ) โ„Ž ( ๐‘ก ) = ๐œ™ ( ๐‘ง ( ๐‘ก ) ) To keep things simple, let us assume ๐œ™ is the identity function, i.e., ๐œ™ ( ๐‘– ) = ๐‘– Consider we have the a final loss ๐ฟ , and computed the derivative of โˆ‚ ๐ฟ โˆ‚ โ„Ž ๐‘‡ for some ๐‘ก = ๐‘‡ Using the update rules, the value of โˆ‚ โ„Ž ๐‘‡ โˆ‚ โ„Ž 1 ย  comes out to be ๐‘ค ( ๐‘Ž ๐‘‡ + ๐‘ ) Main Question: What is the value of a? ย Numerical

Log in for full answers

We've collected over 50,000 authentic original questions and detailed explanations from around the globe. Log in now and get instant access to the answers!

Similar Questions

More Practical Tools for Students Powered by AI Study Helper

Join us and instantly unlock extensive past papers & exclusive solutions to get a head start on your studies!