Content moderation
Purpose and deployment context
Content moderation protects JobsAI from discriminatory, illegal, fraudulent or otherwise harmful content in job advertisements and reports. Currently there is only the rule-based anti-discrimination review of advertisements, application validation rules and a human notice-and-action process. A stand-alone general toxicity classifier is not deployed in the repository or in the production code.
Classification under the AI Act
Under Regulation (EU) 2024/1689 (the AI Act), the current general content moderation is not a stand-alone system for the recruitment evaluation of candidates under Annex III, point 4(a). In its present scope, it is a rule-based and human review of the publication of user content; if it were to begin automatically evaluating candidates or filtering applications, it must be reclassified as high-risk. The Operator nevertheless maintains technical documentation, logging, post-market monitoring and an incident process; serious incidents are reported to the competent authorities under the AI Act and Czech implementing legislation; after the final determination of the supervisory authority, this part will be updated. The obligations for any high-risk extension would take effect on 2 August 2026.
Inputs and outputs
The input is the title and description of a job advertisement, the report text, internal notes and other user-inserted content. These texts may contain contact details, company identifiers and sensitive data if the user inserts them into free text. The output of the current system is a finding of an anti-discrimination rule, a validation error, a human review queue or a DSA notice-and-action record; not a toxicity score from a hosted model.
Models used and versions
As at 2026-05-13, no stand-alone hosted toxicity model was found in the code. The production-documented automated part uses a rule-based anti-discrimination classifier and the ordinary rules of application validation of job advertisements. A general toxicity classifier is planned at the earliest for Q1 after launch; before it is switched on, the model details, DPA/SCC, evaluation set, metrics and a separate approval of the model card must be added. No fine-tuning is recorded. A knowledge cutoff does not apply to the current rule-based part.
Training data
The rule-based part uses an internal evaluation set of anti-discrimination sentences and a legal/policy taxonomy in the documentation. The general toxicity/legality classifier does not yet have a training dataset or prompt in the repository. For human reviews, the DSA reasons, the user's report and the advertisement context are used; the full internal instructions for moderators are not published, so that they cannot easily be circumvented.
Limitations and known risks
The current rule-based review does not catch all forms of fraud, hatred, harassment, illegal offers or toxic content. The card must therefore not be interpreted as a claim that JobsAI operates a general production toxicity classifier. Critical interventions must have human oversight, an appeal and an audit trail. If an external LLM or moderation API is switched on in the future, the quality for Czech and the false-positive impacts must be measured before production use.
Fairness and anti-discrimination
Content moderation takes into account the protected categories under Act No. 198/2009 Coll. and the obligation of non-discriminatory access to employment. The mitigations include the separation of policy reasons, the option of an appeal, the auditing of interventions and a review of impacts by category of breach. The pre-launch audit of the rule-based anti-discrimination part is completed according to the internal methodology. Last audit: 2026-05-14, passed.
| Item | Value |
|---|---|
| Last audit | 2026-05-14, passed |
An audit of the general moderation classifier will take place before it is possibly switched on.
Human oversight
Every hiding or rejection of content must be reviewable. The employer receives the reason and the path of appeal. Trust & Safety can reverse the decision, request an amendment or escalate the incident to the legal contact. DSA notice-and-action, appeals and policy disputes are received at [email protected].
Performance metrics
For the general moderation classifier, precision/recall has not yet been measured, because it is not deployed. It will be measured only before its possible launch: the number of reports, the median time to intervention, the proportion of successful appeals, the false-positive review rate and the category of intervention. The anti-discrimination part has separate metrics in the relevant model card.
Security and privacy
The rule-based part runs locally without transfer to a third party. The processing of personal data is governed by Regulation (EU) 2016/679 (GDPR); the legal basis for moderation, the prevention of misuse and auditing is legitimate interest under Article 6(1)(f) GDPR, and for the statutory notice-and-action obligations also a legal obligation under Article 6(1)(c) GDPR. If an external moderation API is switched on, the texts of advertisements and reports may leave the JobsAI boundary; the Operator must, before launch, add the EU/US processing, the retention regime and DPA/SCC. Moderator queues must mask unnecessary PII and retain the audit trail only for the necessary period. Details are in the privacy policy.
Complaints and contact
Send reports of illegal content, appeals against moderation and complaints about the automated tool to [email protected]. Direct personal data protection requests to the privacy policy.
Card version
Card version 1.0.3, last updated 2026-07-08.