Data Engineer

Chantilly, VA
Full Time
Manager/Supervisor

BT-393 – Data Engineer
Location- Chantilly



**MUST HAVE A TS/SCI CLEARANCE TO APPLY. Those without an active security clearance will not be considered.**





Bespoke Technologies is seeking a Software Developer to provide ETL, Data Engineering, and Full-Stack Software Development support

Required Skills:

  • Demonstrates experience designing and maintaining enterprise-grade ETL/ELT pipelines, both batch and real-time.
  • Demonstrates front-end development and implementation skills, using React, Next.js, or similar.
  • Demonstrates back-end development using Python, Java, Scala, and microservices architecture.
  • Demonstrates experience with API design.
  • Demonstrates experience with containerization, using Docker.
  • Demonstrates experience with CI/CD pipelines.
  • Demonstrates experience with infrastructure-as-code patterns.
  • Demonstrates experience with probabilistic models, risk scoring, Bayesian inference, Monte Carlo simulation, and probabilistic graphical models.
  • Demonstrates experience applying statistical modeling tools.
  • Demonstrates experience with designing cloud-native architectures using cloud services such as AWS, Google, IBM, and Oracle
  • Demonstrates experience designing and operating big data systems
  • Demonstrates experience building and optimizing performance of large scale graph databases (tens of billions of edges) using DynamoDB or new enhanced capabilities
  • Demonstrates experience developing and operating graph traversal capabilities using data graphing tool traversal capabilities built upon Apache Gremlin or new enhanced capabilities
  • Demonstrates experience developing and operating NoSQL solutions to complex big data applications
  • Demonstrates experience in data modeling for performance, partition sharding, record/event aggregation workflows, stream processing, and metrics gathering
  • Demonstrates experience designing and operating large-scale serverless geospatial indexes built with GeoMESA
  • Demonstrates experience with partition and sort key design and implementation to ensure consistent performance
  • Demonstrates experience with aggregation operations to de-duplicate records on continuous data feeds
  • Demonstrates subject matter expertise experience with relational databases to noSQL
  • Demonstrates experience building and operating high performance data processing pipelines using Lambda, Step Functions and PySpark
  • Demonstrates experience building high quality User Interface/User experiences with the React framework and webGL
  • Demonstrates experience designing and operating large scale graph databases using Apache Cassandra
  • Demonstrates experience performing in-depth technical analysis of large-scale graph databases to develop implementation strategies for search optimizations
  • Demonstrates experience developing technical capabilities for processing, persistence and search of datasets that are collected or maintained using standards common in the Sponsor's community
  • Demonstrates experience facilitating engineering discussions across teams representing multiple stakeholders to develop and execute implementation strategies that meet mission needs
  • Demonstrates experience developing Machine Learning Operations (MLOps) pipelines for large scale application
  • Demonstrates experience maintaining configuration of software using configuration management resources such as GitHub
  • Demonstrates experience designing, building and operating big data systems, such as persistence, partitioning, indexing, at scale of trillions of records/events
  • Demonstrates experience with Niagara Files (NiFi) applications or new enhanced capabilities
  • Demonstrates experience developing and operating Kubernetes infrastructure
  • Demonstrates experience supporting engineering efforts that will contribute to delivery of capabilities such as datasets and functionality such as communications, geospatial workflows
  • Demonstrates experience implementing DevSecOps and agile development in production environments
  • Demonstrates experience with agile software development and testing
  • Demonstrates experience with federal security, regulatory and compliance requirements and security accreditation package development
  • Demonstrates experience with data security and governance using centralized security controls like LDAP, encrypting the data, and auditing access to the data
  • Demonstrates experience with specialized technologies that are optimized for the particular use of the data, such as relational databases, a NoSQL database (Cassandra), or object storage
  • Demonstrates experience with Apache, TINKERPOP, GREMLIN and/or JANUSGRAPH to design, develop, implement and maintain system
  • Demonstrates knowledge of Graph Database to design, develop, implement and maintain system
  • Demonstrates experience with C or C++ to write interfaces
  • Demonstrates experience using centralized security controls like LDAP, encrypting data, and auditing access to data
  • Demonstrates experience with:
  • Databases: Postgres, MariaDB, ELK, Minio, AWS S3, Neo4j, MongoDB, noSQL
  • Languages: Python (pypi libraries)
  • Operating Systems: Centos7, RockyLinux8
  • Orchestration: Kubernetes, Docker, Docker-Compose, Docker-Swarm
  • Development Tools: vscode, gitlab, jupyterhub/notebooks, MATLAB
  • Environments: large collaboration and development environments
  • Data types: Unstructured, structured, or semi-structured data, including: CSV, JSON, JSONL, AVRO, Protocol Buffers, Parquet, etc

Desired Skills:
  • Demonstrates experience with designing cloud-native architectures using Sponsors cloud services
  • Demonstrates experience designing and operating big data systems within the Sponsors policy and regulatory environment
  • Demonstrates experience developing and operating graph traversal capabilities using the Sponsors data graphing tool traversal capabilities built upon Apache Gremlin
  • Demonstrates experience building and operating high performance data processing pipelines using Lambda, Step Functions and PySpark on the Sponsors infrastructure with EMR
  • Demonstrates experience working with the Sponsor's enterprise services used for Data Management, including the enterprise catalog service (and associated APIs), and Policy Decision Points (PDPs).
  • Demonstrates experience developing Machine Learning Operations (MLOps) pipelines for large scale application in the Sponsor's environment
  • Demonstrates experience and understanding of IT Service Management and common SLA measurements
  • Demonstrates experience presenting solutions, requirements, and presentations to diverse audiences.
  • Demonstrates experience working with container orchestration technologies such as AWS ECS, AWS Fargate, and Kubernetes or other enhanced capabilities available
  • Demonstrates experience in managing large operational cloud environments spanning multiple tenants using Multi-Account management, AWS Well Architected Best Practices, and AWS Organization Units/Service Control Policies (OU/SCP).
  • Demonstrates experience with Micro-services such as building decoupled systems, utilizing RESTful endpoints and lightweight systems
  • Demonstrates experience in total systems perspectives, including a technical understanding of systems and applications relationships, dependencies, and requirements of hardware and software components
  • Demonstrates experience consulting with customers to determine present and future user needs
  • Demonstrates experience providing frequent contact with customers, traceability within program documents, and the overall computing environment and architecture
 
Desired Certifications:
  • AWS Certified Solutions Architect
  • AWS Machine Learning Certification(s)
  • Agile certification
  • Azure
  • Security+
  • GSEC
  • CCNA
Share

Apply for this position

Required*
We've received your resume. Click here to update it.
Attach resume as .pdf, .doc, .docx, .odt, .txt, or .rtf (limit 5MB) or Paste resume

Paste your resume here or Attach resume file

Human Check*