Tanakh
PulseAugur coverage of Tanakh — every cluster mentioning Tanakh across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
OCR for Hebrew text struggles with vowel points due to preprocessing and training data limitations
Optical Character Recognition (OCR) for Hebrew text often struggles with vowel points (niqqud) due to several preprocessing stages that remove these delicate marks before the recognition model even sees the text. These …
-
New algorithm uses information theory to find formulaic text clusters
Researchers have developed a novel information-theoretic algorithm to identify formulaic clusters within textual data. This method utilizes weighted self-information distributions, extending classical measures to a cont…
-
MiqraBERT model enhances Biblical Hebrew parallel detection
Researchers have developed MiqraBERT, a new Sentence-BERT model specifically finetuned for detecting semantic similarity in Biblical Hebrew. This model, built upon AlephBERT, uses a regression-based approach with cosine…