The data spans structured records, documents, images, audio and video.
You’ll work across bulk ingestion and on-demand retrieval, making this information searchable and useful in our productsYou’ll own systems from source acquisition through to serving queries.
The work includes:Large-scale ingestion.
Build and operate high-throughput, resumable pipelines for large datasets, with efficient incremental updates, monitoring and recovery from failures.
Document processing and data quality.
Extract useful content from complex documents and other formats.
Handle malformed records and changing schemas, and validate outputs while preserving structure and metadata.
Search and serving.
Build keyword, vector and structured search, and design schemas, indexes and partitioning for fast queries over tens to hundreds of millions of records.
Connecting information across sources.
Link patents, scientific records and supporting documents, preserve dates and versions, and make results traceable to their original sources.
Profile parsing, ingestion, database builds and queries throughout development, testing against representative datasets at realistic scale.
Diagnose CPU, memory and storage I/O bottleneck.