Silu Activation Function
PulseAugur coverage of Silu Activation Function — every cluster mentioning Silu Activation Function across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New method extracts bias-free GLU blocks from language models
Researchers have developed a new method for cryptanalytically extracting bias-free Gated Linear Unit (GLU) feed-forward blocks from language models. This technique, which uses finite-difference curvature and paired obse…
-
llama.cpp releases include server improvements and performance optimizations · 8 sources tracked
The llama.cpp project has released several updates, including version b10331 which improves server functionality by correctly reporting the isolate working directory. Other recent releases, such as b10330 and earlier, h…
-
New Kazry activation function outperforms SiLU in neural network tests
A new piecewise activation function for neural networks, named Kazry, has been introduced. Developed by Katriel Fishel Tchursh, Kazry has demonstrated superior performance compared to the SiLU activation function in tests.
-
New framework empirically compresses deep neural networks via state analysis
Researchers have developed a novel method for compressing deep neural networks by analyzing the controllability and observability of their internal states. This framework treats trained networks as dynamical systems, us…
-
New structural interpretation of GELU and other activation functions proposed
Researchers have proposed a new structural interpretation of activation functions like GELU, ReLU, SiLU/Swish, and hard swish. This work views GELU not just as a stochastic gate output, but through a Gaussian complement…
-
New neural network architectures tackle complex scientific computing problems · 8 sources tracked
Researchers are developing novel neural network architectures to solve complex partial differential equations (PDEs) and model dynamical systems. These include structure-oriented randomized neural networks (SO-RaNN) for…
-
New framework enables spiking neural networks for large language models
Researchers have developed a new framework to make large language models more compatible with neuromorphic hardware. The method focuses on creating spike-friendly approximations for the nonlinear operators within Transf…
-
Neural networks achieve super-fast convergence and represent complex functions with floating-point arithmetic
Two new arXiv papers explore theoretical aspects of neural network convergence and representation capabilities. The first paper demonstrates that neural network classifiers can achieve super-fast convergence rates under…