Enter a job title or keyword

Robot Manipulation Video Annotator

Welocalize


Job Location:

Manila - Philippines

Monthly Salary: Not provided by the employer
Posted: 29 September 2026 (2 days ago)
Application Deadline: 27 December 2026
Vacancies: 1 Vacancy

Job Summary

About the Role

We are looking for detail-oriented annotators to help label robot manipulation videos for AI training purposes. Youll watch short videos of robots performing manipulation tasks (filmed from three synchronized camera angles) and produce precise structured natural-language descriptions of the actions taking place. This work directly supports the development of robotics AI models and requires strong written English sharp observational skills and the discipline to follow a detailed style guide consistently.

Project Details

  • Job Title: Project Cursa - Robot Manipulation Video Annotator
  • Location: Remote (Philippines)
  • Work type: Freelance
  • Language: English
  • Pay Rate: US$3.5 per hour

What Youll Do

  • Watch short robot manipulation videos each filmed from three synchronized camera views (an overhead view and views from each of the robots two wrist-mounted cameras).
  • Break each video into time segments and write clear natural-language descriptions for each segment.
  • Apply labels at three levels of detail for each applicable segment:
    • Atomic motion (a few seconds) a single small movement (e.g. close fingers around the red handle)
    • Skill / subtask (several seconds to 20 seconds) a complete meaningful action (e.g. pick up the red block by its edge)
    • Task / goal (up to 1 minute) the overall purpose of a sequence of skills (e.g. place all blocks in the container)
  • Ensure every moment of video is covered by a label at two or more of these levels no gaps including idle or pause moments.
  • Accurately describe exactly what happens including when something doesnt go as planned (a dropped object a failed grasp a slipped grip). Precision matters more than making the robot look successful.
  • Cross-reference all three camera angles: use the overhead view to understand the overall scene and object identity and the close-up wrist views to confirm exact contact and grasp details.
  • Follow a detailed style guide covering vocabulary for actions spatial relationships object descriptions and manner of movement applying it consistently across many episodes.
  • Participate in periodic calibration sessions to align your labeling with the team and the clients reference examples.

What Were Looking For

Required:

  • Strong written English youll write dozens of short precise descriptive sentences per video and need to vary your language rather than repeating the same phrases.
  • Sharp attention to detail able to distinguish small differences (a successful grasp vs. a fumble a push vs. a drag which specific object part is being touched).
  • Comfort following a detailed structured style guide and applying it consistently even in ambiguous or edge-case scenarios.
  • Basic comfort with spatial/mechanical description (left/right above/below naming object parts like handles lids or edges).
  • Reliable self-directed work habits this is often heads-down work with periodic check-ins rather than close supervision.

Nice to Have:

  • Prior experience with video annotation data labeling transcription or QA work.
  • Familiarity with robotics terminology (grippers end-effectors manipulation) helpful but not necessary as the style guide is self-contained.
  • Experience with annotation tools such as Label Studio.

About Company

Company Logo

Welocalize enables brands to reach and grow global audiences through services and solutions for translation, localization, adaptation, interpretation, and automation. We offer multilingual solutions to transform all content types for local audiences, at every step of our clients’ glob ... View more

View Profile View Profile