When an artificial intelligence system learns to tell a cat from a dog, it needs someone to label thousands of pictures first. That plain-sounding job is called data labeling, and China is now pushing it onto a much larger stage.
At a big data industry fair in Guiyang, the National Data Administration said it will start a second round of pilot work in thirty-two cities. Seven cities, including Chengdu and Hefei, already run such pilots, so the total will reach thirty-nine.
The work covers fields such as industrial manufacturing, health care, transport and embodied intelligence, which means machines that can move and act in the physical world. Labelers mark raw data so that it can be used to train AI models.
The supply of data is growing fast. By August this year, China had built more than 126,000 high-quality data sets with a total volume of over 1,815 petabytes, up more than 89 percent from the end of the first quarter. A national data set service system launched in April has already published over 1,700 sets.
Officials also said they will study new business models for tokens, the small units that AI systems use to measure and price their work. Universities will be encouraged to train more data specialists. Labeling may look like low-level work, but without it, even the smartest model has nothing to learn from.
Words to know:
- label (标注,贴标签)
- artificial intelligence (人工智能)
- pilot (试点,试验性的)
- raw data (原始数据)
- petabyte (拍字节)
- token (词元)
- model (模型)
- specialist (专业人员)
长难句拆解 · 过去完成时 + up 引导的增幅表达
例句:By August this year, China had built more than 126,000 high-quality data sets with a total volume of over 1,815 petabytes, up more than 89 percent from the end of the first quarter.
句首的 By August this year 是时间状语,与后面的过去完成时 had built 搭配,表示「到今年 8 月为止已经建成」。英语里凡是出现 by + 过去时间点,谓语通常要用过去完成时,这是判断题和语法题的高频考点。
with a total volume of over 1,815 petabytes 是 with 复合结构,补充说明这些数据集的体量。逗号后面的 up more than 89 percent 是过去分词短语作状语,up 在这里不是「向上」的动作,而是「增长到、增加了」的意思,后面直接跟百分比;from the end of the first quarter 说明比较的基准。中文习惯说「比第一季度末增长超过 89%」,英文则把增幅前移、基准后置,语序正好相反。
直译:到今年 8 月,中国已建成超过 12.6 万个高质量数据集,总体量超过 1815 拍字节,比第一季度末增长超过 89%。
再看一句:The work covers fields such as industrial manufacturing, health care, transport and embodied intelligence, which means machines that can move and act in the physical world.
主干是 The work covers fields,后面的 such as 列举了四个领域。难点在句尾的 which:它引导非限制性定语从句,但指代的不是紧挨着的 embodied intelligence,而是前面整个短语 embodied intelligence ——也就是对「具身智能」这个术语作解释。判断依据是谓语 means 的主语若只是单个名词,语义会变得不通;只有把被解释的概念当作主语才讲得通。这种「which 解释前文术语」的写法在科普文章里极为常见。
直译:这项工作覆盖工业制造、医疗健康、交通运输以及具身智能等领域,具身智能指的是能在物理世界中移动和行动的机器。
同义替换 · 雅思阅读利器
下面 7 组替换都出自本文,右侧是对应的替换说法,点开核对。
- is now pushing it onto a much larger stage → is now taking it to a far bigger level(push onto a larger stage 与 take to a bigger level 同义)
- said it will start a second round of pilot work in thirty-two cities → announced a new wave of trial projects across 32 cities(second round 与 new wave 同义)
- so the total will reach thirty-nine → bringing the overall figure to 39(reach 与 bring to 同义)
- which means machines that can move and act in the physical world → referring to robots that can move and act in real settings(the physical world 与 real settings 同义)
- The supply of data is growing fast → Data provision is expanding rapidly(supply 与 provision 同义)
- up more than 89 percent from the end of the first quarter → an increase of over 89 percent since the close of the first quarter(up … percent 与 an increase of … percent 同义)
- Labeling may look like low-level work → Labeling might appear to be basic work(look like 与 appear to be 同义)
阅读自测 · 3 题
三道题分别对应数字、定义与结论,先不看答案。
-
- 简答:How many cities will take part in data labeling pilot work after the new round?
-
- 判断正误(T / F / NG):Embodied intelligence refers to machines that can move and act in the physical world.
-
- 判断正误(T / F / NG):The passage says data labeling will soon be replaced by AI itself.
参考答案与解析
每题给出答案与判断依据,注意最后提到的干扰项套路。
-
- Thirty-nine. 依据在第二段。原文写 Seven cities … already run such pilots, so the total will reach thirty-nine,问的是新一轮之后的城市总数,答案就是三十九个。答 thirty-two 只算了新增城市,漏掉了已有的七个,属于「以偏概全」。
-
- T(正确)。第三段在列举领域之后紧接 which means machines that can move and act in the physical world,正是对 embodied intelligence 的解释,与题干一致。题干把原文的同位解释改写成了直接定义,属于典型的同义改写。
-
- NG(未提及)。全文说的是数据标注岗位将扩展到三十九个城市、人才需求在扩大,从未提到它会被人工智能取代。题干凭常识补充了一个合理但原文没有的结论,属于「无中生有」——看起来再合理,也只能选 NG。
中文译文
当人工智能系统要学会分辨猫和狗时,它首先需要有人给成千上万张图片打上标签。这份听起来很朴素的工作叫做数据标注,而中国正在把它推向一个更大的舞台。
在贵阳举办的一场大数据产业博览会上,国家数据局表示将在三十二个城市启动第二批先行先试工作。包括成都、合肥在内的七个城市此前已开展此类试点,因此总数将达到三十九个。
这项工作覆盖工业制造、医疗健康、交通运输以及具身智能等领域——所谓具身智能,指的是能够在物理世界中移动和行动的机器。标注员为原始数据打上标记,使其可以用来训练人工智能模型。
数据供给正在快速增长。截至今年 8 月,中国已建成超过 12.6 万个高质量数据集,总体量超过 1815 拍字节,较第一季度末增长超过 89%。今年 4 月上线的一个国家数据集管理服务系统,已发布数据集超过 1700 个。
相关负责人还表示,将研究词元的新型商业模式——词元是人工智能系统用来计量和定价其工作量的最小单位。高校将被鼓励培养更多数据专业人才。标注看起来像是低层次的工作,但没有它,再聪明的模型也无从学习。