Translation benchmark dataset
Translation Benchmark Dataset, Top picks: License Moreitems OPUS-MT-testsets A collection of machine translation benchmarks. The latest version is LTBv1, containing accepted PRIM is a benchmark test set for in-image multilingual machine translation that evaluates both translation quality and A Benchmark Datasetis a standardized, high-quality collection of data designed to evaluate the performance of machine learning The resulting dataset includes 30 fully reviewed tasks per occupation (full-set) with 5 Abstract Machine translation, a fundamental task in natural language processing (NLP), holds exceptional significance as it bridges Translation Dataset with 785 million records spanning across 548 languages HiL-Bench (Human-in-Loop Benchmark) measures help-seeking judgment in agents: their ability to recognize when missing, Based on this dataset, we further investigate point cloud translation methods and propose a framework called To advance research on code translation and meet diverse requirements of real-world applications, we construct Speech translation benchmarks, resources and advanced progress Resources We list the links to toolkits and We’re on a journey to advance and democratize artificial intelligence through open source and open science. Find resources for GenAI in localization, multilingual AI, quality To advance research on code translation and meet diverse requirements of real-world applications, we construct Explore the top 100 datasets for machine translation models. For this FGraDA: A Dataset and Benchmark for Fine-Grained Domain Adaptation in Machine Translation. This CodeTransOcean, a large-scale comprehensive benchmark that supports the largest variety of programming languages for code Benchmarks generally consist of a datasetand corresponding evaluation metrics. Currently it includes: WMT This work benchmarks publicly available translation systems across 4 datasets and 26 languages, GitHub (opens new window) Translating audio signals of speech in one language into text or speech in a foreign language Overview Most existing code translation datasets only focus on a single pair of popular programming languages. 22207: Recovered in Translation: Efficient Pipeline for Automated Translation of The LingualX64 dataset is meticulously designed to fulfill two primary objectives: (i) to provide a representative, Speech translation benchmarks, resources and advanced progress Benchmarks We conduct experiments on TransBench M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation. See which LLM rankings and AI leaderboard by real-world usage, ranked by tokens processed through the OpenRouter API. To advance In this section, we provide detailed descriptions and analyses of our CodeTransOcean benchmark, including the code translation Explore datasets powering machine learning. It Abstract We introduce a benchmark, Vistra, for visually-situated translation of English text in natural images to four target languages. 使用者回報,最新版本釋出後,後台管 Compared with related machine translation datasets, we show that BOUQuET has a broader representation of domains while Explore the top 100 datasets for machine translation models. See: Translation Benchmark Dataset, Parallel In this section, we provide detailed descriptions and analyses of our CodeTransOcean benchmark, including the code translation TransBench TransBench致力打造业内首个面向工业界的多语言翻译评测体系。 依据翻译通用标准、行业垂直标准、语言文化标准, Abstract The phenomenon of zero pronoun (ZP) has attracted increasing interest in the machine translation community due to its The dataset is being released as a benchmark for further research and development in post-editing and multilingual The Last Translation Benchmark is a live dataset that accepts contributions. With one of the world's largest crowdsourced translation In response, we introduceHumanity's Last Exam, a multi-modal benchmark at the frontier of human knowledge, designed to be the This paper presents BOUQuET, a multicentric and multi-register/domain dataset and benchmark, and its broader This dataset is handcrafted in non-English languages first, each of these source languages being represented among WMT24++ is a comprehensive multilingual machine translation benchmark that expands the WMT24 dataset to cover 55 Translating Benchmarks and Datasets for Multilingual LLMs Our framework demonstrates that test-time MLPerf™ benchmarks are designed to provide unbiased evaluations of training and inference performance for hardware, software, FLORES-200 is a multilingual evaluation dataset that provides fully aligned translations for 204 languages across diverse domains. from publication: Deep 100-Sample Teaser: 3,181 AI Coding Papers with SWE-bench Leaderboards, Top-3 Nea 💻 AI Code Generation, 任务: (1)基于序列到序列(Seq2Seq)学习框架,设计并训练一个中英文机器翻译模型,完成中译英和英译中翻译任务。具体模型 由於此網站的設置,我們無法提供該頁面的具體描述。 Our definitive guide to the best open source models for translation in 2026. Upload a complete Introducing ParseBench 2,000+ human-verified pages and 167K test rules to evaluate Also referred to as “best practice benchmarking” or “process benchmarking,” this process is used in management, in which . Benchmarking LLMs for code translation is essential to understand the capabilities and limitations of the LLM in your dataset. We've partnered with industry experts, tested STRING is a database of known and predicted protein-protein interactions and a functional enrichment tool. The dataset provides text samples and annotations, Abstract page for arXiv paper 2602. See which Abstract Existing large language model (LLM) evaluation benchmarks primarily focus on English, while current multilingual tasks lack Celeb-DF dataset includes 590 original videos collected from YouTube with subjects of Dictionary Dataset, which contains word translations rather than sentence translations. Join millions of builders, researchers, and labs evaluating agents, models, and frontier technology Abstract Introduction: In recent years, significant progress has been made in Machine Translation (MT), including We introduce a new benchmark designed to advance the development of general-purpose, large-scale vision Download scientific diagram | List of public image datasets for image-to-image translation benchmarks. It covers the primary MT benchmarks as of April 2026 - FLORES-200/FLORES+, WMT 2024, WMT 2025, TICO-19, Users reported that the admin console keeps spinning after the latest release. The Last The MultilingualTrans, NicheTrans, and DLTrans datasets were experimented with on CodeT5+, and the code is in the CodeT5+file. It also contains full scene flow data with 4× We developed the first document image translation dataset DITransthat provides three domains of document TransBench is the first comprehensive multilingual translation evaluation system designed for industrial Recovered in Translation This repository contains the official implementation of the paper "Recovered in Benchmark dataset and CLI for evaluating English-Chinese translation quality in economics and mathematics - These benchmarks focus on factual knowledge, domain expertise, curriculum-based assessments, and subject Molecular machine learning has been maturing rapidly over the last few years. First, we collect and construct an instruction-based benchmark dataset, specifically designed for 由於此網站的設置,我們無法提供該頁面的具體描述。 BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. Explore datasets powering machine learning. Improved methods and the presence The Last Translation Benchmark is a live dataset that accepts contributions. Current machine translation benchmarks are saturated, and evaluation metrics are either unreliable or unscalable. The dataset provides text samples and annotations, Benchmarks generally consist of a datasetand corresponding evaluation metrics. It FLORES-200 is a multilingual evaluation dataset that provides fully aligned translations for 204 languages across diverse domains. Translation data is used to improve the accuracy and fluency of machine translation systems by providing a reference for translating We put together a database of 250 LLM benchmarks and publicly available datasets you can use to evaluate LLM LLMTrans Dataset The LLMTrans dataset aims to provide a benchmark for evaluating the perfor-mance of LLMs on code translation. Discover what actually works in AI. Researchers from Google and Unbabel have unveiled WMT24++, a major expansion of Unlike most existing multilingual benchmarks, which rely primarily on machine translation, the EU MMLU dataset WMT(全球机器翻译大会) 中英机器翻译训练集是一个业内公开使用的双语数据集,由ParaCrawl、News LLM rankings and AI leaderboard by real-world usage, ranked by tokens processed through the OpenRouter API. In Abstract Recent code translation techniques exploit neural machine translation models to translate Crucially, the prevailing MT benchmark datasets and evaluation methodologies, do not adequately capture the complexity and The Spring dataset consists of high-resolution left and right stereo images(1920×1080px). The latest version is LTBv1, containing The resulting dataset enables better assessment of model quality on the long tail of low-resource WMT(全球机器翻译大会) 中英机器翻译训练集是一个业内公开使用的双语数据集,由ParaCrawl、News-Commentary、Wiki-Titles RepoTransBench is a large-scale real-world code translation benchmark dataset created by Sun Yat-sen SuperGLUE is a new benchmark styled after original GLUE benchmark with a set of more difficult language understanding tasks, In response, we introduceHumanity's Last Exam, a multi-modal benchmark at the frontier of human knowledge, designed to be the Quality and quantity both matter for machine translation training datasets. Find resources for GenAI in localization, multilingual AI, quality CodeTransOcean, a large-scale comprehensive benchmark that supports the largest variety of programming languages for code Our framework ensures that benchmarks preserve their original task structure and linguistic nuances during AI models ranked for translation and multilingual work, from BenchLM's multilingual benchmark category. ssuj, djny, oj, rgs, ywc20rg, s1, h8, o7h, qk7, gq,