How to Measure the Quality of MT Systems: Integrating Evaluation into the Enterprise

Throughout this series of articles we have seen that evaluating a machine translation (MT) system properly involves making a number of decisions: which datasets to use, which automatic metrics to choose and how to interpret their results, when to turn to human evaluation and with what methodology, and how to check that the results are reliable.

At this point, one more practical question remains: how do we integrate all of this into a real workflow?

In this last article of the series, we review some of the tools that exist to make machine translation evaluation easier. We also introduce Vera (González et al., 2026), our MT evaluation platform, which brings automatic evaluation, human evaluation, and comparative analysis of results together in a single environment.

 

The problem: fragmented workflows

There are tools that address different aspects of machine translation evaluation, but each tends to focus on one specific part of the process. MATEO (Vanroy et al., 2023)¹, for example, focuses on calculating automatic metrics such as BLEU, chrF, TER, COMET or BLEURT through a web interface. MT-LENS (Gilabert et al., 2025)² extends the analysis to aspects such as bias or robustness, while Pearmut (Zouhar and Kocmi, 2026)³ is geared towards human evaluation using frameworks such as MQM, DA or ESA, which we looked at in more detail in the article on human evaluation.

In practice, this can mean working with different tools, formats and workflows depending on the type of evaluation you want to carry out. Vera grew out of precisely this need to integrate these different approaches into a single environment and, in addition, it allows you to analyse the relationship between the results of automatic metrics and human judgement.

 

Automatic evaluation

Vera calculates some of the most widely used reference-based metrics: BLEU, chrF++ and TER, through SacreBLEU, and COMET.

To compare systems more rigorously, the platform also makes it possible to analyse the statistical significance of the differences between models using bootstrap resampling for each metric. This way, we can not only tell which system achieves a higher score, but also determine whether the observed difference is statistically significant.

As we saw in the article on automatic evaluation, checking statistical significance is one of the best practices to bear in mind when comparing translation systems.

 

Human evaluation with MQM Core

For human evaluation, Vera uses MQM Core, a subset of the MQM framework that is widely used in the industry.

The platform allows several annotators to work in parallel on the same project, which makes it easier to calculate the agreement between them (Inter-Annotator Agreement, IAA). This measure makes it possible to analyse the extent to which evaluators agree in their annotations and, therefore, provides information about the consistency of the results.

In addition, to make the process more efficient, Vera incorporates sampling strategies (Zouhar et al., 2025). Users can specify what proportion of the corpus they want to evaluate, and the platform automatically selects the most representative and informative segments. In this way, it is possible to reduce the annotation effort without having to evaluate the whole corpus manually, while still retaining a sample that is adequate for obtaining relevant information about system performance.

 

Creating and editing references

Reference-based evaluation requires human reference translations. For this reason, Vera also lets you work with your own reference corpora from the same interface.

Existing references can be edited directly and, where no reference is available, a new one can be created from the output of a machine translation system and its subsequent post-editing.

This makes it possible to prepare and review the data needed for evaluation without having to move it between different tools.

 

Results, correlations and reports

Once the evaluations have been completed, Vera lets you view the results of both automatic and human evaluation within the platform itself and compare the performance of the different systems.

The platform also allows you to analyse the correlation between automatic metrics and human evaluation, at both segment and system level, using Pearson and Spearman coefficients. This analysis makes it possible to study the extent to which automatic metrics correspond to human judgement in a given context.

This is precisely where meta-evaluation comes in: it is not just a matter of knowing which system gets a higher score, but of understanding which metrics provide information closest to human evaluation and in which situations other forms of evaluation may be needed.

Finally, Vera can generate PDF reports and export results, references and annotated corpora in CSV and JSON formats, so you can use them in other analysis workflows.

 

Conclusions

Throughout this series we have covered the main elements involved in machine translation evaluation: from building datasets to automatic metrics and human evaluation.

As we have seen, none of these approaches on its own provides all the information needed to evaluate an MT system. The choice of data, the evaluation methods and the way the results are interpreted are all part of a single process.

Vera brings these different stages together in a single workflow, allowing you to combine automatic and human evaluation, compare systems and analyse the relationship between the results obtained by both approaches.

If you would like to find out more about the platform and see how it works, you can request a demo of Vera. 

 


NOTES

¹ Source: https://huggingface.co/spaces/BramVanroy/mateo-demo 

² Source: https://github.com/langtech-bsc/mt-evaluation 

³ Source: https://github.com/zouharvi/pearmut 

 


BIBLIOGRAPHY

Vanroy, Bram, Arda Tezcan, and Lieve Macken. "MATEO: MAchine translation evaluation online." Proceedings of the 24th Annual Conference of the European Association for Machine Translation. 2023.

Gilabert, Javier García, et al. "MT-LENS: An all-in-one toolkit for better machine translation evaluation." Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations). 2025.

Zouhar, Vilém, and Tom Kocmi. "Pearmut: Human evaluation of translation made trivial." arXiv preprint arXiv:2601.02933 (2026).

Zouhar, Vilém, Peng Cui, and Mrinmaya Sachan. "How to Select Datapoints for Efficient Human Evaluation of NLG Models?". Transactions of the Association for Computational Linguistics 13 (2025): 1789-1811.

González, Sofía García, et al. "VERA: A Platform for Automatic and Human Evaluation of Machine Translation." Proceedings of the 26th Annual Conference of the European Association for Machine Translation (Volume 2). 2026. 

Do you have a project?

Request a no-obligation quote.