このフォルダの語彙ファイルは、次の上流の tokenizer.json を形式だけ変換したものです。 変換は scripts/tokenizer/build_vocab.py(語片・点数・合成の順・正規化の表を抜き出し、使われない正規化の項目を省き、Qwen はバイナリに詰め直した)。 語彙の中身は変えていません。上流はどちらも Apache License 2.0 で、本文は LICENSE-Apache-2.0.txt にあります。 The vocabulary files in this directory are format conversions of the upstream tokenizer.json files listed below (converted by scripts/tokenizer/build_vocab.py; the vocabulary itself is unchanged). Both are licensed under the Apache License, Version 2.0; see LICENSE-Apache-2.0.txt. t5.json Source: google-t5/t5-base, tokenizer.json https://huggingface.co/google-t5/t5-base (revision a9723ea7f1b39c1eae772870f3b547bf6ef7e6c1) Copyright Google LLC. Licensed under the Apache License, Version 2.0. Modified: converted to a compact JSON (pieces, scores, the reachable entries of the precompiled normalization charsmap). qwen.bin Source: Qwen/Qwen3.5-0.8B, tokenizer.json https://huggingface.co/Qwen/Qwen3.5-0.8B (revision 2fc06364715b967f1860aea9cf38778875588b17) Copyright 2026 Alibaba Cloud. Licensed under the Apache License, Version 2.0. Modified: converted to a compact binary (token bytes, merge split points, added tokens).