Price: 0
Number of applications: 11
30.09.26 (inclusive)
Development contract; payment by stages, after written acceptance of each stage. The cost is based on the contractor's commercial offer. Long-term cooperation in the subcontracting format on Xena Soft projects is possible.
MVP
ICT tasks
Media sphere
Neurotechnology and artificial Intelligence
Software/ IS
Xena Soft is an IT service company that develops web products for customers from the United States and Europe. In each project, E2E tests become obsolete faster than the team manages to maintain them. What we see on the projects: - The frontend is deployed several times per hour, and a full run of E2E tests takes from one to three hours. The queue of runs is growing, and the tests are checking the stand, which has since been re-deployed several more times. - No one analyzes a dropped run: most often the reason is a changed selector, not a product, and the team stops trusting unsuccessful runs. - The coverage is not growing. New features and bug fixes are released without regression tests, because there is no one to write tests. - Some of the products (the web build of the React Native application, native components) do not have semantic markup, and tests cannot reliably find the elements. - There is no dedicated test automation engineer for each project: the tests are supported by developers in their free time from developing new features. A working prototype showed that an AI agent based on Claude Code copes with this: it corrects tests, writes new ones, creates defect tasks, and conducts cumulative PR/MR. But it is impossible to transfer the prototype to the next project without rewriting it: it is tied to a specific CI system and branch scheme.
- For failures that the agent can sort out: it takes no more than 60 minutes from the publication of artifacts of the dropped run to the PR/MR with correction, excluding the CI queue. The actual indicator is measured on the pilot. - Every unsuccessful run gets into a classification story. The team receives separate notifications about alleged product defects and cases that the agent has failed to handle; the agent corrects outdated selectors without the participation of the team. - Coverage grows without the participation of a QA engineer: after a successful run, the agent opens a PR/MR only if there is a confirmed lack of coverage. Guideline: Every bug fix in the product receives a regression test within a week. - It takes no more than one working day to connect a project that meets the preliminary requirements from the documentation. The creation and verification of the basic test suite are evaluated separately.
Ostapova Xenia Sergeevna
Purpose and description of task (project)
To develop a portable AI agent that runs inside the continuous integration (CI) pipeline and supports end-to-end automated tests (E2E tests) on Playwright for web products. All code changes are processed by the agent as merge requests (PR/MR); they are accepted only after a review. After each deployment, the agent runs the tests on the stand and classifies the result.: 1. Tests have failed — analyzes the log, traces, and screenshots and determines the source of the failure: the test, product, or infrastructure. The outdated test corrects and opens the PR/MR. The alleged product defect is framed as a task with supporting materials and an indication of the likely location of the defect in the code. 2. Tests passed — compares coverage with product changes made after the last processed commit. If he finds an uncovered change, he writes a script, runs it on the stand and opens PR/MR. If there are no such changes, it fixes "no changes required". 3. There are no tests — it creates a set from scratch according to the approved list of scenarios. If the interface elements cannot be reliably identified, first prepares a separate PR/MR with stable selectors, and writes the tests in the next run after deployment. 4. If the reason cannot be determined, the booth is unavailable or there is not enough data, it completes the run with the status `blocked`, describes what needs to be clarified, and does not change the code. Xena Soft independently developed a working prototype and uses it on two of its projects: one pipeline on Azure DevOps Pipelines, the other on GitHub Actions. Each project has its own conventions, and configuration, access rights, and branches are configured separately. It is necessary to turn the prototype into a product: a single core, three CI adapters (GitHub Actions, GitLab CI, Azure DevOps), a panel with a history of runs and metrics, documentation according to which a new project is connected in one working day.
Note
The candidate sends a brief description of the architectural approach, a preliminary estimate of the time and cost, the composition of the team and an impersonal example of similar work. After the initial selection, the parties sign the NDA, and the selected performer receives an anonymized copy of the prototype. The work proceeds in stages; each stage ends with a demonstration and written acceptance. A pilot for one Xena Soft project lasts 4-6 weeks; the criteria for its acceptance are given in the terms of reference