loader icon
AI & Data Engineering Full-time India

Data Engineer — AI & Spatial Data Infrastructure

Bhubaneswar, India (SPARC Global Office)

Role overview

You will build and maintain the data infrastructure backbone for SPARC's AI and geospatial product suite — including TraceOS (document intelligence), Atlas (geospatial AI), and the Terrascope platform (MineScope, AgriScope, ForestScope).

You will build ETL pipelines ingesting diverse data sources, ensure data quality through parsing and validation, enable AI/ML workflows through embeddings and vector databases, and maintain data governance with audit logs and lineage tracking.

Key responsibilities

  • Build connectors for diverse data sources: SharePoint, Box, Teams, SFTP, S3, APIs, databases, and IoT platforms.
  • Handle multiple formats: PDF, Excel, Word, CAD, GeoTIFF, shapefiles, GeoJSON, satellite imagery, and sensor streams.
  • Build incremental ingestion pipelines with error handling, retry logic, and dead-letter queues for failed ingestions.
  • Preserve document structure during parsing — tables, captions, headers, footers, and figure references.
  • Extract and normalize units, coordinate systems, and temporal formats; extract metadata including dates, authors, and spatial bounds.
  • Process satellite imagery and GIS layers — raster tiling, cloud masking, vector spatial indexing, and CRS transformations.
  • Integrate spatial data infrastructure with PostGIS and geospatial databases.
  • Implement automated data quality checks — schema validation, unit verification, spatial/temporal bounds checking, and de-duplication.
  • Chunk documents and spatial features for semantic search; generate and store embeddings in vector databases with rich metadata.
  • Implement audit trails, data lineage tracking, and policy enforcement for retention, access control, and compliance.
  • Build and maintain quality dashboards and alerting for data quality degradation.

Required skills

  • 3+ years of Python in production environments, with strong command of pandas, NumPy, and asyncio.
  • Experience building data pipelines with Apache Airflow, Databricks, Azure Data Factory, or Snowflake, processing 10,000+ files per day reliably.
  • Strong experience with PostgreSQL for structured metadata and at least one vector database (Pinecone, Weaviate, or Qdrant).
  • Hands-on experience with document parsing libraries — PyPDF2, pdfplumber, Camelot, openpyxl, python-docx, and OCR tools (Tesseract, Azure Document Intelligence, or AWS Textract).
  • Strong understanding of cloud deployment (Azure preferred, or AWS), Docker containerization, and CI/CD pipelines for data workflows.
  • Cost-conscious engineering mindset — avoiding unnecessary API calls and optimizing storage.
  • Proficiency with AI coding assistants (Cursor, Claude, GitHub Copilot) to accelerate development and debugging.

Nice to have

  • Experience with Databricks or Snowflake for large-scale data processing and warehousing.
  • Graph databases (Neo4j) for data relationships and lineage tracking.
  • Advanced geospatial tooling — GDAL, rasterio, Fiona, and satellite imagery processing.
  • CAD parsing (ezdxf for DXF) or LiDAR data handling (LAS/LAZ).
  • API development (FastAPI, Flask) for data services and quality dashboards.
  • Stream processing experience (Kafka, Kinesis) for real-time data pipelines.
  • Experience with compliance or regulatory data in environmental, mining, or agriculture sectors.
  • Familiarity with STAC (SpatioTemporal Asset Catalog) for geospatial metadata.

Why join SPARC Global

  • Work on production AI infrastructure that powers compliance-grade systems for environmental firms and mining operators.
  • Direct ownership of ETL architecture, tool choices, and scaling decisions.
  • Master document intelligence, vector databases, and compliance-grade data quality.
  • Work closely with Parsa (AI Lead), ML engineers, and backend engineers across Bhubaneswar and Canada.
  • Competitive compensation and benefits package.