Data Entry Jobs in Lebanon
1825 Jobs Found
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p><b><br></b></p><p><b>About the Role </b>Are you passionate about building modern, scalable data platforms that power business decisions and product innovation? As a Staff Data Engineer at Toters, you'll define the technical direction of our data platform while designing and building reliable, high-performance data infrastructure that scales with our business. You'll lead the architecture of our batch and streaming data ecosystem, establish engineering standards, and drive platform evolution across multiple teams. Working closely with Engineering Managers, Software Engineers, Product Managers, Analytics, and Machine Learning teams, you'll influence technical strategy, mentor engineers, and ensure our data platform remains scalable, reliable, and cost-efficient. This is a highly hands-on technical leadership role where you'll shape how data is ingested, processed, governed, and consumed across the organization.</p><p>In this role, you will:</p><ul><li>Define and drive the long-term technical vision and architecture for Toters' data platform.</li><li>Lead the design and implementation of scalable batch and streaming data platforms that support analytics, operational systems, and machine learning.</li><li>Design resilient, cloud-native data architectures capable of processing high-volume event streams and transactional workloads.</li><li>Establish engineering standards, architectural patterns, and best practices for data engineering across the organization.</li><li>Design scalable data models, schemas, and data contracts that enable long-term maintainability, governance, and interoperability.</li><li>Drive architectural decisions that improve scalability, reliability, security, observability, and operational efficiency across the platform.</li><li>Lead complex technical initiatives such as platform migrations, streaming modernization, lakehouse evolution, and data infrastructure improvements.</li><li>Evaluate new technologies and recommend pragmatic solutions that balance scalability, operational complexity, and cost.</li><li>Improve platform observability through monitoring, alerting, data quality validation, lineage, and operational dashboards.</li><li>Own the reliability of critical production data systems by leading incident response, root cause analysis, and long-term platform improvements.</li><li>Optimize cloud infrastructure, storage, compute, and streaming costs while maintaining high platform performance.</li><li>Partner with Engineering Managers, Product Managers, Analytics, Machine Learning, and Software Engineering teams to define platform capabilities and technical roadmaps.</li><li>Mentor Senior Data Engineers through architecture reviews, technical coaching, and engineering leadership.</li><li>Conduct high-quality code and design reviews that raise the technical bar across multiple teams.</li><li>Drive technical discussions, resolve complex engineering challenges, and influence cross-functional decision-making across the engineering organization.</li></ul><p>Why Toters?</p><ul><li>Flexible work environment with hybrid-friendly roles.</li><li>Opportunity to define and shape the future of Toters' modern data platform.</li><li>Lead high-impact technical initiatives that influence multiple engineering teams.</li><li>Solve complex engineering challenges involving distributed systems, streaming, and large-scale data processing.</li><li>Collaborate with talented engineers in a culture of mentorship, ownership, and continuous learning.</li><li>Direct impact on products used by thousands of customers every day.</li><li>Competitive compensation package.</li><li>Exclusive discounts on Toters orders.</li><li>First-class medical insurance.</li></ul></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><ul><li>Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.</li><li>9+ years of experience designing and building large-scale production data platforms or distributed data systems.</li><li>Deep expertise in SQL and advanced data modeling, including dimensional modeling (facts and dimensions), schema design, and data architecture.</li><li>Expert-level experience building distributed streaming systems using Apache Kafka, Amazon Kinesis, Apache Pulsar, or equivalent technologies, with strong knowledge of partitioning, consumer groups, delivery semantics, schema management, and production troubleshooting.</li><li>Deep understanding of distributed systems, batch and streaming data processing, event time, processing time, windowing, stateful processing, fault tolerance, and the challenges of achieving exactly-once processing.</li><li>Strong programming skills in Python and/or JVM languages such as Java or Scala.</li><li>Extensive experience designing modern data lakehouse architectures using Amazon S3, Parquet, and open table formats such as Apache Iceberg, Delta Lake, or Apache Hudi.</li><li>Experience building cloud-native data platforms on AWS, Google Cloud Platform (GCP), or Microsoft Azure, including object storage, streaming services, IAM, networking, infrastructure security, and cloud cost optimization.</li><li>Hands-on experience with distributed processing technologies such as Apache Flink, ClickHouse, Trino, and similar analytical platforms.</li><li>Experience designing and implementing Change Data Capture (CDC) architectures using Debezium or similar technologies.</li><li>Experience with modern transformation frameworks such as dbt.</li><li>Strong experience with Infrastructure as Code using Terraform, containerization using Docker, and orchestration platforms such as Kubernetes.</li><li>Proven experience owning mission-critical production systems, leading incident response, and driving long-term platform reliability.</li><li>Demonstrated ability to define engineering standards, influence architectural decisions, and lead technical initiatives across multiple teams.</li><li>Excellent communication and technical writing skills, with the ability to document architecture, influence stakeholders, and mentor senior engineers.</li><li>Passion for building scalable, reliable, maintainable, and cost-efficient data platforms.</li><li>Nice to Have Experience designing enterprise-scale streaming architectures using Apache Flink.</li><li>Experience leading lakehouse migrations using Apache Iceberg or similar technologies.</li><li>Experience implementing large-scale CDC platforms using Debezium.</li><li>Experience building self-service data platforms and developer tooling.</li><li>Experience with ML Feature Stores, Customer Data Platforms (CDPs), or real-time personalization platforms.</li><li>Experience working in high-growth technology companies, marketplace platforms, or on-demand delivery businesses.</li><li>Experience contributing to open-source technologies or internal engineering frameworks.</li></ul><p></p></section>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>Job Overview The Microsoft Fabric Data Engineer designs, builds, and operates modern data platforms using Microsoft Fabric. This role focuses on ingesting, modeling, and serving data via OneLake, Lakehouse, Data Warehouse, Data Pipelines, and Power BI delivering trusted, performant datasets and governed analytics at scale. The role collaborates closely with data architects, analytics engineers, BI developers, and business stakeholders. Key Responsibilities 1. Data Platform Engineering (Fabric) Build and manage Lakehouses (Delta Lake) and Fabric Data Warehouses. Develop Data Pipelines and Dataflows Gen2 for batch and near-real-time ingestion. Create and optimize Notebook-based transformations (PySpark/SQL) and SQL stored procedures for DW workloads. Implement medallion architecture (bronze/silver/gold) for scalable curation. Publish certified semantic models and Power BI datasets aligned to business domains. 2. Performance & Reliability Optimize storage/compute in OneLake (file formats, partitioning, z-ordering). Tune Spark and SQL workloads (caching strategies, concurrency, workload isolation). Implement robust retry, alerting, and monitoring (Fabric Monitoring Hub, Metrics app). Conduct end-to-end pipeline performance testing and scalability assessments. 3. Governance, Security & Compliance Enforce data governance with sensitivity labels, row-level/column-level security, and workspace roles. Manage item-level permissions (Lakehouse tables, DW schemas, datasets) and Managed Identities for sources. Apply data quality rules, lineage, and documentation (Descriptions, Tags, Owner metadata; Purview if applicable). Ensure compliance with organizational standards (PII handling, audit, retention). 4. DevOps & Lifecycle Management Use Fabric Git integration and Deployment Pipelines for CI/CD across dev/test/prod. Parameterize pipelines and environments; externalize configuration and secrets (Key Vault). Implement automated testing for data transformations and schemas. Drive release management, change control, and rollback strategies. 5. Collaboration & Stakeholder Engagement Partner with analytics engineers and BI teams to design star schemas, semantic models, and DAX measures. Work with data source owners for SLAs, schema change management, and contracts. Translate business requirements into technical designs and document architecture decisions. Provide knowledge transfer, best practices, and support to data consumers.</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><h2>Required Skills & Qualifications</h2><ul><li><strong>Technical Skills</strong> Microsoft Fabric (hands-on): OneLake, Lakehouse (Delta), Fabric Data Warehouse, Data Pipelines, Dataflows Gen2, Notebooks, Semantic Models/Power BI, Monitoring Hub.</li><li><strong>Programming & Querying</strong>: PySpark, SQL (T-SQL), Delta Lake operations; DAX familiarity is a plus.</li><li><strong>Modeling & Architecture</strong>: Dimensional modeling, Data Vault or medallion patterns, data quality frameworks.</li><li><strong>Performance & Ops</strong>: Partitioning, file formats (Parquet/Delta), caching/z-ordering, job orchestration, monitoring.</li><li><strong>DevOps</strong>: Git, Fabric Deployment Pipelines, YAML CI/CD (GitHub Actions/Azure DevOps), IaC exposure (Bicep/Terraform for non-Fabric infra).</li><li><strong>Security & Governance</strong>: RLS/CLS, sensitivity labels, access patterns, audit/logging, lineage.</li><li><strong>Preferred Qualifications</strong> Experience with Power BI modeling (star schemas, relationships, calculation groups, DAX). Exposure to streaming/real-time: Eventstream, Real-Time Hub, KQL databases (if applicable). Experience integrating with external sources (SQL Server, SAP, Dataverse, REST APIs). Familiarity with Microsoft Purview for governance/lineage.</li><li><strong>Certifications</strong>: DP-600: Microsoft Fabric Analytics Engineer Associate (strongly preferred) DP-203: Data Engineering on Microsoft Azure (nice to have)</li></ul><p></p></section>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>We are seeking a Data Pipeline & Warehousing Engineer to design, build, and maintain reliable data pipelines and data warehouses in Lebanon, enabling scalable analytics and reporting for business and technical stakeholders. Job Purpose Own end-to-end ETL/ELT and data warehousing solutions ensuring data quality, high performance, and governed data models while automating monitoring, alerting, and operational support for analytics workloads. Job Duties and Responsibilities SQL (T-SQL/PostgreSQL/MySQL) ETL/ELT (Airflow, dbt) Python Cloud platforms (AWS/GCP/Azure) Big data tools (Spark) Data warehousing (Snowflake/BigQuery/Redshift) Version control (Git) Monitoring/logging (CloudWatch/Stackdriver) Problem-solving and analytical thinking Communication with stakeholders</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><ul><li>SQL (T-SQL/PostgreSQL/MySQL)</li><li>ETL/ELT (Airflow, dbt)</li><li>Python</li><li>Cloud platforms (AWS/GCP/Azure)</li><li>Big data tools (Spark)</li><li>Data warehousing (Snowflake/BigQuery/Redshift)</li><li>Monitoring/logging (CloudWatch/Stackdriver)</li><li>Version control (Git)</li><li>Data governance / governed data models</li><li>Strong analytical and troubleshooting skills</li><li>Stakeholder communication skills</li></ul><p></p></section>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
Company Description<br>
Jobs for Humanity is partnering with hassan+721 to build an inclusive and just employment ecosystem. Therefore, we prioritize individuals coming from all walks of life.
<br> Company Name: hassan+721<br>
<br>Job Description<br>
<p>We are seeking a Data Pipeline & Warehousing Engineer to design, build, and maintain reliable data pipelines and data warehouses in Lebanon, enabling scalable analytics and reporting for business and technical stakeholders.</p><br> Job Purpose <p>Own end-to-end ETL/ELT and data warehousing solutions—ensuring data quality, high performance, and governed data models—while automating monitoring, alerting, and operational support for analytics workloads.</p><br> Job Duties and Responsibilities <ul>
<li>SQL (T-SQL/PostgreSQL/MySQL)</li><li>ETL/ELT (Airflow, dbt)</li><li>Python</li><li>Cloud platforms (AWS/GCP/Azure)</li><li>Big data tools (Spark)</li><li>Data warehousing (Snowflake/BigQuery/Redshift)</li><li>Version control (Git)</li><li>Monitoring/logging (CloudWatch/Stackdriver)</li><li>Problem-solving and analytical thinking</li><li>Communication with stakeholders</li>
</ul>
<br>
<br>Qualifications<br>
Required Qualifications <ul>
<li>SQL (T-SQL/PostgreSQL/MySQL)</li><li>ETL/ELT (Airflow, dbt)</li><li>Python</li><li>Cloud platforms (AWS/GCP/Azure)</li><li>Big data tools (Spark)</li><li>Data warehousing (Snowflake/BigQuery/Redshift)</li><li>Monitoring/logging (CloudWatch/Stackdriver)</li><li>Version control (Git)</li><li>Data governance / governed data models</li><li>Strong analytical and troubleshooting skills</li><li>Stakeholder communication skills</li>
</ul>
<br>
<br><br> </div>
We are looking for an experienced and detail-oriented Data Analyst to join our team in Dora. The successful candidate will be responsible for analyzing business data, preparing reports, identifying trends, and providing actionable insights to support better decision-making across the company.
<br>
<br>Key Responsibilities
<br>Collect, clean, validate, and analyze data from various sources.
<br>Prepare accurate and timely reports, dashboards, and performance analyses.
<br>Identify trends, patterns, and opportunities through data analysis.
<br>Monitor key performance indicators (KPIs) and provide regular updates to management.
<br>Work closely with different departments to understand their data and reporting needs.
<br>Translate business requirements into meaningful analytical reports and insights.
<br>Ensure data accuracy, consistency, and integrity.
<br>Automate and improve existing reporting and data-analysis processes.
<br>Present findings and recommendations clearly to both technical and non-technical stakeholders.
<br>Requirements
<br>Bachelor’s degree in Data Analytics, Business, Statistics, Computer Science, Mathematics, Economics, or a related field.
<br>2+ years of proven experience as a Data Analyst or in a similar analytical role.
<br>Strong proficiency in Microsoft Excel, including PivotTables, formulas, and data analysis.
<br>Experience with SQL and databases.
<br>Experience with data visualization tools such as Power BI, Tableau, or similar is an advantage.
<br>Strong analytical and problem-solving skills.
<br>Excellent attention to detail and accuracy.
<br>Ability to communicate complex data clearly and effectively.
<br>Ability to manage multiple priorities and meet deadlines.
<br>Fluency in English; Arabic is a plus
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>Job Overview The Data Quality Engineer will ensure the accuracy, consistency, and reliability of data across the Microsoft Fabric Data Warehouse, its integrations, and connected services. The role focuses on testing and validating data pipelines, event-driven processes, and Azure-based integrations such as Logic Apps and Service Bus. The engineer will collaborate closely with data engineers, architects, and business units to maintain high-quality, trustworthy data for analytics and reporting. Key Responsibilities Design, develop, and execute data quality and validation tests for Microsoft Fabric Data Warehouse objects (tables, views, semantic models). Validate ETL/ELT pipelines, including Fabric Dataflows and integrated source systems, for accuracy and completeness. Test and monitor event-driven data processes, including Azure Service Bus messages, queues, and topics. Validate Azure Logic Apps workflows and integrations that move or transform data between systems. Perform end-to-end testing of data ingestion, transformations, and delivery to downstream consumers or BI systems. Develop and maintain automated data quality checks, reconciliation scripts, and test frameworks. Identify, analyze, and report data quality issues, including root cause analysis and impact assessment. Collaborate with Data Engineers, BI teams, and business units to define data quality rules, metrics, and acceptance criteria. Monitor data quality KPIs, implement controls to prevent defects, and ensure reliability in production pipelines. Document test cases, workflows, data quality rules, and testing outcomes clearly for technical and business stakeholders. Familiarity with unit testing, automated data validation, and integration testing tools such as dbt, Great Expectations, tSQLt, pytest, Azure Logic Apps testing, and Service Bus testing scripts. Perform unit testing and validation of Azure Logic Apps workflows and Service Bus message flows. Develop and maintain automated data quality checks and integration tests using tools such as dbt, Great Expectations, or custom scripts. Validate and test API endpoints and integration workflows using tools such as Postman to ensure data is accurately transmitted between systems. Perform end-to-end testing of event-driven processes and service integrations using Postman and automated scripts.</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><p>We are looking for entrepreneurs, techies, passionate, eager to learn, humble, with a positive attitude and a high level of integrity People. Flexible and willing to take challenges, work and live in coffee-producing countries, People who want to build expertise and a career in the coffee business and are ready to go the extra mile.</p><p>Required Skills & Qualifications</p><ul><li>Strong experience in data quality testing and validation for data warehouses or large-scale analytics platforms.</li><li>Hands-on experience with Microsoft Fabric (Data Warehouse, Lakehouse, Pipelines, Dataflows).</li><li>Proficiency in SQL and experience validating ETL/ELT pipelines.</li><li>Experience with event-driven architectures and Azure integrations: Azure Service Bus (queues, topics, subscriptions) Azure Logic Apps for workflow automation</li><li>Experience testing data integrations across multiple source systems.</li><li>Familiarity with data quality concepts: completeness, accuracy, consistency, timeliness, uniqueness.</li><li>Experience with automation and scripting (Python, PySpark, PowerShell).</li><li>Understanding of data modeling concepts: star schema, snowflake, fact and dimension tables.</li><li>Strong analytical, problem-solving, and troubleshooting skills.</li><li>Excellent documentation and communication skills.</li><li>Familiarity with unit testing, automated data validation, and integration testing tools such as dbt, Great Expectations, tSQLt, pytest, Postman, Azure Logic Apps testing, and Service Bus testing scripts .</li></ul><p>Preferred / Nice to Have:</p><ul><li>Experience with data quality frameworks (e.g., Great Expectations).</li><li>Knowledge of Azure Data Factory, Synapse Analytics, and OneLake.</li><li>Familiarity with CI/CD pipelines for data platforms.</li><li>Understanding of data governance, metadata management, and auditing standards.</li></ul><p></p></section>
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p><br></p><p><b>About the Role</b> Are you passionate about building scalable, reliable data platforms that power business decisions and customer experiences? As a Senior II Data Engineer at Toters, you'll design and build modern data infrastructure that enables reliable, real-time and batch data processing across the organization. You'll own the design and delivery of complex data pipelines, influence technical decisions within your domain, and help establish engineering best practices that improve the scalability, reliability, and maintainability of our data platform. Working closely with Software Engineers, Product Managers, Analytics, and Machine Learning teams, you'll help transform high-volume operational data into trusted, high-quality data products that drive business impact.</p><p>In this role, you will:</p><ul><li>Collaborate with cross-functional teams to understand business requirements and translate them into scalable data solutions.</li><li>Design, build, and maintain reliable batch and streaming data pipelines that support analytics, operational systems, and machine learning workloads.</li><li>Develop scalable data ingestion frameworks capable of processing high-volume event and transactional data.</li><li>Design and optimize data models that enable efficient analytics while supporting long-term maintainability.</li><li>Build resilient streaming solutions while balancing throughput, latency, reliability, and operational cost.</li><li>Develop and maintain modern data lakehouse architectures using cloud-native technologies and distributed data platforms.</li><li>Design schemas and data contracts that promote consistency, governance, and interoperability across systems.</li><li>Improve platform observability by implementing monitoring, alerting, data quality validation, and operational dashboards.</li><li>Troubleshoot production issues, participate in incident response, perform root cause analysis, and implement long-term improvements.</li><li>Optimize storage, compute, and data movement costs while maintaining platform performance and reliability.</li><li>Participate in architecture discussions and contribute to technical decisions that improve scalability and system resilience.</li><li>Create reusable frameworks, tooling, and engineering patterns that improve developer productivity and platform consistency.</li><li>Conduct thoughtful code reviews and share constructive feedback to maintain high engineering standards.</li><li>Mentor junior engineers through technical guidance, design discussions, and knowledge sharing.</li><li>Collaborate with Software Engineering, Product, Analytics, and Machine Learning teams to deliver scalable, high-quality data solutions.</li></ul><p>Key Qualifications</p><p>Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.</p><p>6+ years of experience designing and building production-grade data platforms or distributed data systems.</p><p>Strong SQL skills and deep understanding of data modeling, including dimensional modeling (facts and dimensions).</p><p>Hands-on experience building production data pipelines using distributed streaming platforms such as Apache Kafka, Amazon Kinesis, or Apache Pulsar, with experience in partitioning, consumer groups, schema management, delivery semantics, and troubleshooting production streaming systems.</p><p>Strong understanding of batch and streaming data processing concepts, including event time, processing time, windowing, stateful processing, and the trade-offs of exactly-once processing.</p><p>Proficiency in Python and/or JVM languages such as Java or Scala for building scalable data processing applications.</p><p>Experience designing and operating modern data lake or lakehouse architectures using Amazon S3 or equivalent object storage, Parquet, and open table formats such as Apache Iceberg, Delta Lake, or Apache Hudi.</p><p>Experience building cloud-native data infrastructure on AWS, Google Cloud Platform (GCP), or Microsoft Azure, including object storage, streaming and messaging services, IAM, networking, and cost optimization.</p><p>Experience working with distributed processing frameworks such as Apache Flink or similar streaming technologies.</p><p>Familiarity with analytical query engines such as ClickHouse or Trino.</p><p>Experience implementing Change Data Capture (CDC) solutions using Debezium or similar technologies.</p><p>Experience with modern data transformation frameworks such as dbt is a plus.</p><p>Familiarity with Infrastructure as Code using Terraform, containerization with Docker, and orchestration platforms such as Kubernetes is a plus.</p><p>Experience owning production systems, participating in on-call rotations, and resolving production data incidents.</p><p>Strong communication skills with the ability to document technical decisions, collaborate across teams, and mentor engineers.</p><p>Passion for building reliable, scalable, and maintainable data platforms.</p><p>Nice to Have</p><p>Experience building large-scale streaming architectures using Apache Flink.</p><p>Hands-on experience with Apache Iceberg, ClickHouse, or Trino in production environments.</p><p>Experience implementing CDC pipelines using Debezium or similar technologies.</p><p>Experience using dbt for transformation and analytics engineering workflows.</p><p>Experience with AWS data services such as MSK, S3, Glue, or Kinesis.</p><p>Experience deploying infrastructure using Terraform and Kubernetes.</p><p>Exposure to Machine Learning feature stores or Customer Data Platforms (CDPs).</p><p>Experience working in high-growth technology companies or on-demand marketplace platforms.</p><p>Why Toters?</p><p>Flexible work environment with hybrid-friendly roles.</p><p>Opportunity to work on a modern, cloud-native data platform that powers products used by thousands of customers every day.</p><p>Collaborate with talented engineers while growing your technical expertise and career.</p><p>Strong culture of mentorship, collaboration, and continuous learning.</p><p>Opportunity to solve complex engineering challenges involving streaming, distributed systems, and large-scale data processing.</p><p>Direct impact on business-critical products and engineering decisions.</p><p>Competitive compensation package.</p><p>Exclusive discounts on Toters orders.</p><p>First-class medical insurance.</p></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><ul><li>Bachelor's degree in Computer Science, Engineering, Information Systems, or a related field.</li><li>6+ years of experience designing and building production-grade data platforms or distributed data systems.</li><li>Strong SQL skills and deep understanding of data modeling, including dimensional modeling (facts and dimensions).</li><li>Hands-on experience building production data pipelines using distributed streaming platforms such as Apache Kafka, Amazon Kinesis, or Apache Pulsar, with experience in partitioning, consumer groups, schema management, delivery semantics, and troubleshooting production streaming systems.</li><li>Strong understanding of batch and streaming data processing concepts, including event time, processing time, windowing, stateful processing, and the trade-offs of exactly-once processing.</li><li>Proficiency in Python and/or JVM languages such as Java or Scala for building scalable data processing applications.</li><li>Experience designing and operating modern data lake or lakehouse architectures using Amazon S3 or equivalent object storage, Parquet, and open table formats such as Apache Iceberg, Delta Lake, or Apache Hudi.</li><li>Experience building cloud-native data infrastructure on AWS, Google Cloud Platform (GCP), or Microsoft Azure, including object storage, streaming and messaging services, IAM, networking, and cost optimization.</li><li>Experience working with distributed processing frameworks such as Apache Flink or similar streaming technologies.</li><li>Familiarity with analytical query engines such as ClickHouse or Trino.</li><li>Experience implementing Change Data Capture (CDC) solutions using Debezium or similar technologies.</li><li>Experience with modern data transformation frameworks such as dbt is a plus.</li><li>Familiarity with Infrastructure as Code using Terraform, containerization with Docker, and orchestration platforms such as Kubernetes is a plus.</li><li>Experience owning production systems, participating in on-call rotations, and resolving production data incidents.</li><li>Strong communication skills with the ability to document technical decisions, collaborate across teams, and mentor engineers.</li><li>Passion for building reliable, scalable, and maintainable data platforms.</li><li>Experience building large-scale streaming architectures using Apache Flink.</li><li>Hands-on experience with Apache Iceberg, ClickHouse, or Trino in production environments.</li><li>Experience implementing CDC pipelines using Debezium or similar technologies.</li><li>Experience using dbt for transformation and analytics engineering workflows.</li><li>Experience with AWS data services such as MSK, S3, Glue, or Kinesis.</li><li>Experience deploying infrastructure using Terraform and Kubernetes.</li><li>Exposure to Machine Learning feature stores or Customer Data Platforms (CDPs).</li><li>Experience working in high-growth technology companies or on-demand marketplace platforms.</li></ul><p></p></section>
Job Summary
<br>We are looking for a reliable and detail-oriented Data Operations Officer to join our team. The ideal candidate will be responsible for monitoring daily system operations, maintaining accurate data, generating reports, and ensuring data quality. This role requires strong analytical skills and the ability to work accurately in a fast-paced environment.
<br>
<br>
<br>
<br>Key Responsibilities
<br>Monitor daily system activities and ensure data is accurate and up to date.
<br>Review and validate data to identify and correct any discrepancies.
<br>Generate daily and monthly operational reports.
<br>Analyze data and provide insights to support business decisions.
<br>Coordinate with different departments to resolve data-related issues.
<br>Maintain data integrity and follow company data management procedures.
<br>Assist in improving operational processes and reporting efficiency.
<br>Perform other data-related tasks as assigned.
<br>
<br>
<br>Qualifications
<br>Bachelor's degree in Accounting, Auditing, Business Administration, Management Information Systems (MIS), Data Analytics, or a related field.
<br>1–3 years of experience or more in data operations, reporting, accounting, auditing, or a similar role.
<br>Strong knowledge of Microsoft Excel.
<br>Good analytical and problem-solving skills.
<br>High attention to detail and accuracy.
<br>Good communication and organizational skills.
<br>Ability to manage multiple tasks and meet deadlines.
<br>
<br>
<br>How to Apply
<br>If you are detail-oriented, enjoy working with data, and are passionate about improving operational efficiency, we encourage you to apply by submitting your CV & Cover letter
<p><h4>Description</h4>
<p>We're looking for a hands-on data quality analyst to ensure the accuracy, consistency, and reliability of data presented through dashboards, reports, and digital applications.<br>
You'll play a key role in validating business-critical data before it reaches end users, ensuring reports, KPIs, and analytical outputs are accurate, reconciled, and aligned with approved business definitions. Working closely with analytics, BI, product, and engineering teams, you'll help maintain the quality and integrity of data that drives business decisions.</p>
<h4>Requirements</h4>
<h4>What you'll do</h4>
<ul>
<li>Validate data displayed across dashboards, reports, web, and mobile applications</li>
<li>Reconcile source data, transformed datasets, and final analytical outputs</li>
<li>Verify KPIs, metrics, filters, aggregations, and calculation logic used in reports</li>
<li>Ensure consistency of metrics across multiple dashboards and reporting platforms</li>
<li>Validate business rules and ensure reported figures align with approved definitions</li>
<li>Identify data discrepancies and investigate inconsistencies before release</li>
<li>Support user acceptance testing (UAT) for dashboards, reports, chatbots, and other analytical products</li>
<li>Act as the quality gate before dashboards and reports are released to end users</li>
<li>Log, track, and support the resolution of data quality issues</li>
<li>Work closely with BI developers, data engineers, and analytics teams to resolve reporting issues</li>
<li>Maintain documentation for reconciliation rules, report logic, and validation processes</li>
<li>Contribute to improving reporting standards, validation checklists, and quality practices</li>
</ul>
<h4>You're our match if you have</h4>
<ul>
<li>3–5 years of experience in data quality, business intelligence, analytics, or reporting assurance</li>
<li>Hands-on experience validating dashboards, reports, or business intelligence solutions</li>
<li>Strong SQL skills for data validation, reconciliation, and comparison</li>
<li>Experience working with Power BI, Tableau, or similar reporting platforms</li>
<li>Strong understanding of KPIs, metrics, business rules, and calculation logic</li>
<li>Experience working with enterprise data sources such as ERP, CRM, or HCM systems</li>
<li>Excellent analytical skills with exceptional attention to detail</li>
<li>Ability to communicate data issues clearly to both technical and business stakeholders</li>
<li>Strong sense of ownership and commitment to data accuracy</li>
</ul>
<h4>Bonus points if you have experience with</h4>
<ul>
<li>Validating analytics for web or mobile applications</li>
<li>Government or enterprise reporting environments</li>
<li>Data governance, reporting standards, or data quality frameworks</li>
<li>Supporting executive-level reporting</li>
</ul></p><p></p>
Job Summary
<br>We are looking for a reliable and detail-oriented Data Operations Officer to join our team. The ideal candidate will be responsible for monitoring daily system operations, maintaining accurate data, generating reports, and ensuring data quality. This role requires strong analytical skills and the ability to work accurately in a fast-paced environment.
<br>
<br>
<br>
<br>Key Responsibilities
<br>Monitor daily system activities and ensure data is accurate and up to date.
<br>Review and validate data to identify and correct any discrepancies.
<br>Generate daily and monthly operational reports.
<br>Analyze data and provide insights to support business decisions.
<br>Coordinate with different departments to resolve data-related issues.
<br>Maintain data integrity and follow company data management procedures.
<br>Assist in improving operational processes and reporting efficiency.
<br>Perform other data-related tasks as assigned.
<br>
<br>
<br>Qualifications
<br>Bachelor's degree in Accounting, Auditing, Business Administration, Management Information Systems (MIS), Data Analytics, or a related field.
<br>1–3 years of experience or more in data operations, reporting, accounting, auditing, or a similar role.
<br>Strong knowledge of Microsoft Excel.
<br>Good analytical and problem-solving skills.
<br>High attention to detail and accuracy.
<br>Good communication and organizational skills.
<br>Ability to manage multiple tasks and meet deadlines.
<br>
<br>
<br>How to Apply?
<br>If you are detail-oriented, enjoy working with data, and are passionate about improving operational efficiency, we encourage you to apply by submitting your CV & Cover letter
Job Summary
<br>We are looking for a reliable and detail-oriented Data Operations Officer to join our team. The ideal candidate will be responsible for monitoring daily system operations, maintaining accurate data, generating reports, and ensuring data quality. This role requires strong analytical skills and the ability to work accurately in a fast-paced environment.
<br>
<br>
<br>
<br>Key Responsibilities
<br>Monitor daily system activities and ensure data is accurate and up to date.
<br>Review and validate data to identify and correct any discrepancies.
<br>Generate daily and monthly operational reports.
<br>Analyze data and provide insights to support business decisions.
<br>Coordinate with different departments to resolve data-related issues.
<br>Maintain data integrity and follow company data management procedures.
<br>Assist in improving operational processes and reporting efficiency.
<br>Perform other data-related tasks as assigned.
<br>
<br>
<br>Qualifications
<br>Bachelor's degree in Accounting, Auditing, Business Administration, Management Information Systems (MIS), Data Analytics, or a related field.
<br>1–3 years of experience or more in data operations, reporting, accounting, auditing, or a similar role.
<br>Strong knowledge of Microsoft Excel.
<br>Good analytical and problem-solving skills.
<br>High attention to detail and accuracy.
<br>Good communication and organizational skills.
<br>Ability to manage multiple tasks and meet deadlines.
<br>
<br>
<br>How to Apply?
<br>
<br>If you are detail-oriented, enjoy working with data, and are passionate about improving operational efficiency, we encourage you to apply by submitting your CV & Cover letter
<section><p class="heading jdMain">Job Description</p><p class="heading">Roles & Responsibilities</p><div class="paragraph"><p>A confidential AI-driven technology company serving international clients across the UK, Europe, the US, and the GCC/MENA region is hiring a Data Scientist.</p><p>Responsibilities:</p><ul><li>Analyze large datasets to extract actionable business insights</li><li>Build and validate predictive models and statistical analyses</li><li>Develop dashboards and data visualizations for client reporting</li><li>Collaborate with engineering and product teams to integrate data solutions</li><li>Communicate findings clearly to both technical and non-technical stakeholders</li></ul></div></section><section><p class="heading">Desired Candidate Profile</p><p class="paragraph"></p><ul><li>1+ years of experience in data science or analytics</li><li>Proficiency in Python or R and data manipulation libraries (Pandas, NumPy)</li><li>Experience with machine learning frameworks and statistical modeling</li><li>Familiarity with SQL and cloud data platforms</li><li>Strong analytical thinking and communication skills</li><li>English proficiency required</li></ul><p></p></section>
We are seeking a Data Scientist to join our remote team and work on data-driven and analytical projects.
<br>
<br>Please submit your application through:
<br>https://creativetechy.com/careers/
• Analyze large volumes of trading, financial, and operational data.
<br>• Prepare, maintain, and improve Excel-based reports, analysis, and dashboards.
<br>• Monitor trading activity and identify trends, anomalies, discrepancies, and unusual patterns.
<br>• Perform statistical and quantitative analysis to support trading and management decisions.
<br>• Produce daily, weekly, and ad-hoc reports for the Trading Department and management.
<br>• Validate data for accuracy, consistency, and completeness.
<br>• Reconcile and investigate differences between datasets and trading records.
<br>• Develop and improve spreadsheets, formulas, models, and analytical tools.
<br>• Support traders and management with data-driven insights and analysis.
<br>• Help automate repetitive reporting and data-analysis processes where possible.
<br>• Work with multiple data sources and maintain accurate and reliable datasets.
<br>• Assist with performance analysis and other quantitative projects within the trading environment.
<br>• Respond quickly and accurately to ad-hoc analytical requests in a high-volume environment
<p><h4>Description</h4>
<p>We're looking for a hands-on data quality analyst to ensure the accuracy, consistency, and reliability of data presented through dashboards, reports, and digital applications.<br>
You'll play a key role in validating business-critical data before it reaches end users, ensuring reports, KPIs, and analytical outputs are accurate, reconciled, and aligned with approved business definitions. Working closely with analytics, BI, product, and engineering teams, you'll help maintain the quality and integrity of data that drives business decisions.</p>
<h4>Requirements</h4>
<h4>What you'll do</h4>
<ul>
<li>Validate data displayed across dashboards, reports, web, and mobile applications</li>
<li>Reconcile source data, transformed datasets, and final analytical outputs</li>
<li>Verify KPIs, metrics, filters, aggregations, and calculation logic used in reports</li>
<li>Ensure consistency of metrics across multiple dashboards and reporting platforms</li>
<li>Validate business rules and ensure reported figures align with approved definitions</li>
<li>Identify data discrepancies and investigate inconsistencies before release</li>
<li>Support user acceptance testing (UAT) for dashboards, reports, chatbots, and other analytical products</li>
<li>Act as the quality gate before dashboards and reports are released to end users</li>
<li>Log, track, and support the resolution of data quality issues</li>
<li>Work closely with BI developers, data engineers, and analytics teams to resolve reporting issues</li>
<li>Maintain documentation for reconciliation rules, report logic, and validation processes</li>
<li>Contribute to improving reporting standards, validation checklists, and quality practices</li>
</ul>
<h4>You're our match if you have</h4>
<ul>
<li>3–5 years of experience in data quality, business intelligence, analytics, or reporting assurance</li>
<li>Hands-on experience validating dashboards, reports, or business intelligence solutions</li>
<li>Strong SQL skills for data validation, reconciliation, and comparison</li>
<li>Experience working with Power BI, Tableau, or similar reporting platforms</li>
<li>Strong understanding of KPIs, metrics, business rules, and calculation logic</li>
<li>Experience working with enterprise data sources such as ERP, CRM, or HCM systems</li>
<li>Excellent analytical skills with exceptional attention to detail</li>
<li>Ability to communicate data issues clearly to both technical and business stakeholders</li>
<li>Strong sense of ownership and commitment to data accuracy</li>
</ul>
<h4>Bonus points if you have experience with</h4>
<ul>
<li>Validating analytics for web or mobile applications</li>
<li>Government or enterprise reporting environments</li>
<li>Data governance, reporting standards, or data quality frameworks</li>
<li>Supporting executive-level reporting</li>
</ul></p><p></p>
<p><strong>Job Description Role Overview:</strong> The Microsoft Fabric Data Engineer is responsible for designing, building, optimizing, and operating enterprise-scale data platforms on Microsoft Fabric. This role focuses on data ingestion, transformation, storage, governance, automation, and platform reliability using OneLake, Lakehouse, Data Warehouse, Data Pipelines, and Spark technologies. The position is heavily focused on Data Engineering and Platform Engineering rather than reporting and dashboard development. The successful candidate will build scalable, governed, and high-performance data solutions that support analytics, AI, operational reporting, and business intelligence initiatives across the organization. This role is primarily focused on Data Engineering, Data Platform Development, and Microsoft Fabric architecture. Candidates whose experience is primarily centered around Power BI report development, dashboard creation, or data visualization without substantial Data Engineering experience may not be a fit for this position.</p><p><strong>Key Responsibilities</strong></p><ol><li><strong>Data Engineering & Platform Development</strong> Design and implement enterprise-scale Lakehouse architectures using OneLake and Delta Lake. Build and maintain robust batch, incremental, CDC, and near real-time data ingestion pipelines. Develop scalable ETL/ELT solutions using Fabric Data Pipelines, Dataflows Gen2, PySpark, and SQL. Implement and manage Medallion Architecture (Bronze, Silver, Gold). Develop reusable and metadata-driven ingestion and transformation frameworks. Integrate data from ERP systems, SAP, REST APIs, SQL Server, Dataverse, and other enterprise applications. Design and maintain enterprise data models supporting analytical and operational workloads.</li><li><strong>Fabric Engineering & Optimization</strong> Develop and optimize Fabric Lakehouses and Data Warehouses. Build advanced Notebook-based transformations utilizing PySpark and Spark SQL. Implement Delta Lake capabilities including: Merge/Upsert Change Data Feed (CDF) Time Travel Schema Evolution Optimize Vacuum Design scalable storage, partitioning, and file management strategies within OneLake. Optimize Spark workloads and Data Warehouse performance.</li><li><strong>Performance, Reliability & Monitoring</strong> Tune large-scale Spark and SQL workloads. Implement monitoring, alerting, and operational dashboards using Fabric Monitoring Hub and Metrics. Conduct performance testing and scalability assessments. Troubleshoot pipeline failures and platform performance bottlenecks. Define and monitor SLAs for critical data assets and pipelines. Ensure high availability and operational excellence across data workloads.</li><li><strong>DevOps & Platform Automation</strong> Implement CI/CD using Fabric Git Integration and Deployment Pipelines. Build automated deployment processes across Development, Test, and Production environments. Develop automated testing frameworks for data quality, schema validation, and regression testing. Manage environment configurations, secrets, and deployment parameters. Utilize Azure DevOps or GitHub Actions to support release automation and governance. Define rollback, release management, and change control processes.</li><li><strong>Data Quality, Governance & Security</strong> Implement automated data quality validation frameworks. Develop reconciliation, completeness, and consistency checks. Manage data lineage, metadata, and documentation standards. Implement sensitivity labels, access controls, auditing, and compliance requirements. Ensure adherence to organizational standards for governance, security, retention, and PII handling. Manage workspace permissions, Managed Identities, RLS, and CLS where required.</li><li><strong>Technical Leadership & Collaboration</strong> Translate business requirements into scalable technical solutions. Collaborate with Data Architects, Data Analysts, Integration Engineers, and business stakeholders. Lead engineering best practices and platform standards. Mentor junior engineers and support knowledge-sharing initiatives. Produce architecture documentation and technical design specifications. Drive continuous improvement of the data platform and engineering practices.</li></ol><p><strong>Desired Candidate Profile</strong></p><h2>Qualifications and Experience:</h2><ul><li>Bachelor's degree in Computer Science, Data Engineering, Information Systems, or a related field.</li><li>Minimum 5 years of experience in Data Engineering.</li><li>Minimum 2 years of hands-on experience with Microsoft Fabric, Azure Data Engineering, Databricks, or modern Lakehouse platforms.</li><li>Proven experience designing and building enterprise-grade data platforms.</li></ul><h2>Required Technical Skills</h2><ul><li>Microsoft Fabric</li><li>OneLake</li><li>Lakehouse</li><li>Delta Lake</li><li>Fabric Data Warehouse</li><li>Data Pipelines</li><li>Dataflows Gen2</li><li>Notebooks</li><li>Monitoring Hub</li><li>Deployment Pipelines</li><li>Git Integration</li><li>Data Engineering</li><li>PySpark</li><li>Spark SQL</li><li>SQL / T-SQL</li><li>ETL / ELT Development</li><li>Data Modeling</li><li>Incremental Processing</li><li>Change Data Capture (CDC)</li><li>Data Quality Frameworks</li><li>Medallion Architecture</li><li>Performance Optimization</li><li>Delta Lake Optimization</li><li>Partitioning Strategies</li><li>File Management</li><li>Query Optimization</li><li>Spark Performance Tuning</li><li>Workload Management</li><li>Integration</li><li>REST APIs</li><li>SQL Server</li><li>SAP</li><li>Azure Storage</li><li>Dataverse</li><li>Enterprise Application Integration</li><li>DevOps</li><li>Git</li><li>Azure DevOps</li><li>GitHub Actions</li><li>YAML Pipelines</li><li>CI/CD</li><li>Automated Testing</li><li>Security & Governance</li><li>Row-Level Security (RLS)</li><li>Column-Level Security (CLS)</li><li>Sensitivity Labels</li><li>Managed Identities</li><li>Data Lineage</li><li>Audit and Compliance Controls</li></ul>
Key Requirements:
<br>
<br>Proven experience in data analysis and building detailed reports.
<br>Strong proficiency in data visualization and handling large datasets.
<br>Strong analytical and problem-solving skills
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
<span>We're looking for a hands-on Data Quality Analyst to ensure the accuracy, consistency, and reliability of data presented through dashboards, reports, and digital applications.<br> You'll play a key role in validating business-critical data before it reaches end users, ensuring reports, KPIs, and analytical outputs are accurate, reconciled, and aligned with approved business definitions.<br> Working closely with analytics, BI, product, and engineering teams, you'll help maintain the quality and integrity of data that drives business decisions.<br> What You'll Do ❏ Validate data displayed across dashboards, reports, web, and mobile applications ❏ Reconcile source data, transformed datasets, and final analytical outputs ❏ Verify KPIs, metrics, filters, aggregations, and calculation logic used in reports ❏ Ensure consistency of metrics across multiple dashboards and reporting platforms ❏ Validate business rules and ensure reported figures align with approved definitions ❏ Identify data discrepancies and investigate inconsistencies before release ❏ Support User Acceptance Testing (UAT) for dashboards, reports, chatbots, and other analytical products ❏ Act as the quality gate before dashboards and reports are released to end users ❏ Log, track, and support the resolution of data quality issues ❏ Work closely with BI developers, data engineers, and analytics teams to resolve reporting issues ❏ Maintain documentation for reconciliation rules, report logic, and validation processes ❏ Contribute to improving reporting standards, validation checklists, and quality practices You're Our Match If You Have ❏ 3–5 years of experience in Data Quality, Business Intelligence, Analytics, or Reporting Assurance ❏ Hands-on experience validating dashboards, reports, or business intelligence solutions ❏ Strong SQL skills for data validation, reconciliation, and comparison ❏ Experience working with Power BI, Tableau, or similar reporting platforms ❏ Strong understanding of KPIs, metrics, business rules, and calculation logic ❏ Experience working with enterprise data sources such as ERP, CRM, or HCM systems ❏ Excellent analytical skills with exceptional attention to detail ❏ Ability to communicate data issues clearly to both technical and business stakeholders ❏ Strong sense of ownership and commitment to data accuracy Bonus Points If You Have Experience With ❏ Validating analytics for web or mobile applications ❏ Government or enterprise reporting environments ❏ Data governance, reporting standards, or data quality frameworks ❏ Supporting executive-level reporting</span> </div>
<h2 class="h5">Job description</h2>
<div class="t-break" data-jb-field="description">
Result of Service<br>• One Excel dataset files detailing climate finance flows to Arab Region and Arab States with tabs for each indicator and its associated visualization, and a methodology tab detailing classification system applied and related notes. • Three sector-specific Excel dataset files with tabs for each indicator and its associated visualization in each Excel file, and a methodology tab in each file detailing the classification system applied and related notes. • One Excel datasets that overs comparative analysis between international climate finance flows with other data sources. o Updated Excel dataset files (up to five datafiles) based on ESCWA review, as appropriate. o Five PowerPoint slide decks based on the vetted datasets and visualizations provided in Excel data files, with the vetted key messages on each slide.<br> Work Location<br>Remote<br> Expected duration<br>6 weeks<br> Duties and Responsibilities<br>I. Background The United Nations Economic and Social Commission for Western Asia (ESCWA) analysis of climate finance data flows supports its collaboration with the Government of Sweden on Climate Resilience through Regional Cooperation for Inclusive Sustainable Development. This work is conducted under the umbrella of the Arab Centre for Climate Change Policies, which is situated in the Climate Change and Natural Resource Sustainability Cluster at ESCWA and has a dedicated work pillar related to climate finance. II. Objective The data analyst would provide updated data sets regarding climate finance flows to Arab States at the regional and country level, with three sector specific datasets provided with associated visualizations. III. Duties and responsibilities Under the direct supervision of the Cluster Lead for Climate Change and Natural Resources Sustainability, the consultant would: (a) Deliver datasets and analysis of aggregated international climate finance flows to Arab Region and each Arab State (1 data file) and for the water, energy and food sectors (3 data files) for various indicators based on the existing classification systems applied by ESCWA. (b) Develop visualization for each indicator and at least three key messages describing the findings for each indicator. (c) Prepare comparative analysis between international climate finance flows to the Arab Region with three other financial datasets (1 data file). (d) Submit the five dataset files to ESCWA for review and incorporate comments received. (e) Prepare five PowerPoint slide decks that present the finalized visualizations and key messages, based on the five data files.<br> Qualifications/special skills<br>A master’s degree in statistics, mathematics or related field is required. A first-level university degree with two additional years of relevant work experiences may be accepted in lieu of the advanced university degree. Note: All candidates must submit a copy of the required educational degree. Incomplete applications will not be reviewed. A minimum of five years of experience working on data analytics related to climate change, natural resources management, environment or related fields is required. Demonstrated professional experience working on climate finance datasets is required. Professional proficiency using Microsoft Office applications (Word, Excel, PowerPoint) is required.<br> Languages<br>English and French are the working languages of the United Nations Secretariat; and Arabic is a working language of ESCWA. For this position, fluency in English and Arabic is required. Note: “Fluency” equals a rating of ‘fluent’ in all four areas (speak, read, write, and understand).<br> Additional Information<br>Not available.<br> No Fee<br>THE UNITED NATIONS DOES NOT CHARGE A FEE AT ANY STAGE OF THE RECRUITMENT PROCESS (APPLICATION, INTERVIEW MEETING, PROCESSING, OR TRAINING). THE UNITED NATIONS DOES NOT CONCERN ITSELF WITH INFORMATION ON APPLICANTS’ BANK ACCOUNTS.<br><br><br>
</div>
<p><h4>Mindrift is looking for highly skilled senior Python data scraping engineers</h4>
<p>to join the Tendem project and drive specialized data scraping workflows within our hybrid AI + human system.</p>
<p>In this role, as an AI pilot – that’s how we refer to this role at Mindrift – you’ll collaborate with Tendem agents that handle repetitive tasks, while you provide critical thinking, domain expertise, and quality control to deliver accurate and actionable results.</p>
<p>This part-time remote opportunity is ideal for technical professionals with hands-on experience in web scraping, data extraction, and processing.</p>
<h4>What we do</h4>
<p>The Mindrift platform connects specialists with AI projects from major tech innovators. Our mission is to unlock the potential of generative AI by tapping into real-world expertise from across the globe.</p>
<p>This is a freelance role for a Tendem project. As a senior Python data scraping engineer, you'll handle data scraping tasks requiring technical precision for web extraction and processing, utilizing various tools such as our provided Apify and OpenRouter alongside your own resourceful approaches.</p>
<h4>Key responsibilities:</h4>
<li>Own end-to-end data extraction workflows across complex websites, ensuring complete coverage, accuracy, and reliable delivery of structured datasets.</li>
<li>Leverage internal tools (Apify, OpenRouter) alongside custom workflows to accelerate data collection, validation, and task execution while meeting defined requirements.</li>
<li>Ensure reliable extraction from dynamic and interactive web sources, adapting approaches as needed to handle JavaScript-rendered content and changing site behavior.</li>
<li>Enforce data quality standards through validation checks, cross-source consistency controls, adherence to formatting specifications, and systematic verification prior to delivery.</li>
<li>Scale scraping operations for large datasets using efficient batching or parallelization, monitor failures, and maintain stability against minor site structure changes.</li>
<h4>Requirements:</h4>
<li>At least 5+ years of relevant experience in data engineering, web scraping, automation, or software development (required).</li>
<li>Bachelor’s or master’s degree in engineering, applied mathematics, computer science, or related technical fields is a plus.</li>
<li>Candidates should have a strong technical foundation and practical experience with scripting, automation, and AI-assisted workflows. We are looking for specialists who can solve non-trivial problems, work confidently with LLMs, and systematically collect, structure, and validate data from diverse sources. A methodical, detail-oriented approach and the ability to work independently are essential.</li>
<li>Strong experience in Python web scraping (BeautifulSoup, Selenium or similar), including dynamic content (JS, AJAX, infinite scroll) and APIs via proxies.</li>
<li>Proven ability to extract data from complex structures (hierarchies, archived pages, inconsistent HTML).</li>
<li>Solid background in data cleaning, normalization, and validation, delivering structured datasets (CSV, JSON, Google Sheets).</li>
<li>Demonstrated experience handling anti-bot mechanisms and dynamic site structures at scale.</li>
<li>Experience with cloud infrastructure (AWS or equivalent) and containerization (Docker) as part of real workflows.</li>
<li>Hands-on experience with LLM frameworks (LangChain, OpenRouter, or similar) applied to automation tasks.</li>
<li>Strong attention to detail and commitment to data accuracy.</li>
<li>Self-directed work ethic with ability to troubleshoot independently.</li>
<li>A link to GitHub is a plus.</li>
<li>English proficiency: Upper-intermediate (B2) or above (required).</li>
<h4>Project time expectations</h4>
<p>For this project, tasks are estimated to require around 10–20 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active.</p>
<h4>Compensation</h4>
<p>On this project, contributors can earn up to $25 per hour equivalent, depending on their level and pace of contribution.</p>
<p>Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.</p></p><p></p>