ServicesWorkAboutBlog Contact Start a project

Media and advertising

AI for media and advertising.

Agencies, production houses, broadcasters and publishers run on repetition: the same asset cut fourteen ways, the same brief re-briefed, the same approval chain buried in email. We build the systems that absorb that work — and we are direct about which parts of it AI should not touch.


Mumbai is where this industry already is

Film, television, advertising, music and news publishing sit in one metropolitan region, which means the people who commission this work and the people who do it are in the same city.

Production and post runs from Andheri and Goregaon out to Film City; the agency and network offices sit in Andheri, Lower Parel and BKC. Our office is in Chakala, Andheri East, in the middle of it.

Proximity changes the first meeting rather than the invoice. Content operations problems cannot be scoped from a requirements document, because nobody writes down the parts that are merely annoying. They get scoped by sitting next to a producer for two hours and watching which spreadsheet holds the version numbers, which WhatsApp group carries the approvals, and how many times a file is re-exported because someone changed a supering. The rest of what we do is on the Mumbai page; this one is about media, because the failure modes here are different.

Start with content operations, because it pays back fastest

The largest recoverable cost in a media business is rarely production. It is the handling of what has already been produced: versioning, conforming, routing, chasing and re-exporting.

One campaign asset becomes a sixteen-by-nine master, a nine-by-sixteen cut, a one-by-one, three durations, two language variants, a version with the legal line and one without, platform-specific bitrates, and stills pulled from the same footage. Fourteen deliverables from one edit, and in most shops each is produced by a person opening a project file by hand.

Underneath sits a versioning problem nobody owns: a master somewhere, several near-copies, and a naming convention three people follow and two do not. When a client asks why the wrong cut went live, the honest answer is usually that nobody could tell which cut was right. Around both sits an approval chain living in email, where the record of who approved which version exists only in people's memories.

This layer is unglamorous and it is the fastest payback of anything we build in this sector, because it is deterministic work with clear success criteria. Automated conforming driven off a single master. A canonical asset record with real version identity. Approval routing where the approval attaches to a version rather than to a message. Most of it is workflow automation rather than machine learning, which is exactly why it holds: there is no model to be uncertain about. The test we apply is whether the automation removes a step or moves it. If a producer still checks every output, you have not automated anything; you have added a review queue.

Localisation is not translation, and dubbing is not localisation

Translating the words is the easy part. Register and lip-sync decide whether an audience stays.

India is not one localisation market, and the differences between a Tamil, Bengali and Bhojpuri version are not vocabulary. Register is the first problem: Hindi marks formality distinctions English does not, so a line that is neutral in English forces a choice about how one character addresses another that the original never made. Get it wrong and the character changes. No automated quality score catches that.

Length is the second. Indian languages do not compress the way English does, so a line fitting a two-second shot in English may need three in Malayalam. Delivery speeds up, meaning gets trimmed, or the cut changes. Every dubbing pipeline makes that trade-off; most make it silently.

Lip-sync is the hardest. Matching visible mouth shapes to a different phoneme set is a video generation problem, not an audio one, and it degrades predictably: on close-ups, on fast speech, on profile angles, on partly occluded faces. A pipeline convincing on a mid-shot of a presenter can fall apart on two actors talking over each other. We triage content before processing rather than after, and keep a human gate on the shots where failures cluster. More on the video localisation and dubbing page.


Ad creative at volume, and the measurement problem underneath

Generating two hundred variants is straightforward. Knowing which of them worked, and why, is the actual engineering problem — and most volume-creative projects fail on the second half.

The mechanics are genuinely easy now: hundreds of on-brand permutations from a template system and a small set of assets, and ad creative automation is among the more reliable things to build. Then the arithmetic arrives. Split a fixed budget across two hundred variants and each gets a hundredth of the traffic. Most will never accumulate enough conversions to separate their performance from noise, so you end up with a leaderboard that reranks itself on every refresh and a team drawing conclusions from it.

What makes this work is not the generation but the measurement design: deciding what you are testing before generating anything, structuring variants so a difference is attributable to one dimension rather than six at once, setting a minimum volume per cell and refusing to read a cell below it, and reporting results with their own confidence rather than as a bare number. Done that way, volume creative is a real advantage, particularly for the localisation-heavy campaigns common here. Done without it, it is an expensive way to generate the feeling of optimisation.

Rights and provenance: what you can train on, what you label, what you must prove

Three separate questions that get collapsed into one, and only the third is really an engineering problem.

What you can train on is a rights question. An archive is not a training set merely because you hold it: talent agreements written before generative tools existed rarely contemplate a likeness being used to synthesise new performance, and library footage carries terms that vary per clip. That is your counsel's call. What we can do is make it answerable, by building the asset record so every item carries its licence terms and permitted uses as structured data rather than as a PDF in a folder. Most clients who cannot answer "can we train on this" cannot answer it because the information is not in queryable form.

What you must label is where the rules have moved. Under the Information Technology (Intermediary Guidelines and Digital Media Ethics Code) Amendment Rules, 2026, notified by MeitY on 10 February 2026 and in force from 20 February 2026, synthetically generated information covers audio, visual and audio-visual material that appears real and is likely to be perceived as indistinguishable from a real person or event; text-only content and routine editing sit outside the definition. The duties fall on intermediaries: label visual synthetic content prominently, prefix audio with an audio disclosure, embed permanent metadata or a unique identifier, and do not permit that label or metadata to be removed. Significant social media intermediaries must also obtain a declaration from the uploader and verify it automatically. The draft's numeric thresholds were dropped before notification in favour of a qualitative prominence standard, so the percentage figure still circulating in the trade press comes from a superseded document.

For advertising specifically, ASCI's Draft Guidelines for Responsible Labelling of Synthetically Generated Content in Advertising, dated 8 May 2026, propose three tiers: misleading or infringing content stays prohibited whether or not it is labelled; content where the synthetic element could materially influence a consumer decision requires disclosure, using wording such as "Audio/Video created using AI"; and routine colour correction, noise reduction and decorative backgrounds require none. They are draft and self-regulatory, and we would not build a pipeline that assumes their final wording.

What you must prove later is the part we can actually solve, and the part clients underinvest in. If a platform asks you to declare whether an asset is synthetic, someone has to answer, and that answer should be derived rather than remembered. That means provenance captured at generation time and carried through every transform: which tool produced it, which sources went in, which human approved it, and what the declaration said at publication. Build it once and both the platform declaration and any later dispute become a lookup. Skip it and you are reconstructing history from file names. We build to the requirement your legal team sets; we do not tell you what the law obliges you to do, and you should be wary of any agency that does.


"AI writes the copy" is the least valuable part of this

It is the most demonstrable capability and the smallest line in your cost base, which is exactly why it dominates the pitch decks.

Copy generation demos well: something visible in ten seconds, no integration, and anyone in the room can judge it. So it leads the conversation, and projects get built around it. Even ASCI's draft treats the generation of advertising copy as low risk, alongside colour correction and noise reduction — a fair reflection of where the leverage sits.

The reason is arithmetic. Writing is a small share of the hours in a campaign and a smaller share of the cost, and the writing that carries commercial weight — the concept, the strategic line — is what generic output is worst at. Meanwhile the export queue, the versioning, the approval chasing and the conforming consume a large share of a studio's capacity and are almost entirely mechanical. Automating the first is a rounding error. Automating the second changes what the business can take on.

There is a second-order cost too: volume-generated copy raises review burden. Someone has to read it and check it against the claims and the brand guidelines, and at sufficient volume reviewing costs more than writing did. So we open a media engagement by measuring where the hours go rather than by asking which model you want. It is duller and it produces better projects.

What we will not do

Some of this work is technically straightforward and we still decline it.

We do not synthesise a person's voice or likeness without a written grant that specifically contemplates synthetic use. A general talent release is not that. We do not build systems designed to make synthetic content harder to detect, including anything that strips provenance metadata or defeats a platform's labelling. We do not take on catalogue dubbing where the review stage has been cut to make the numbers work, because a pipeline without a human gate is fine most of the time and its failures land in public. And we do not subcontract: the people in the scoping call are the people who build it, which is a real constraint on how much we take on at once.

Our office is at B-204, Kanakia Wall Street, Andheri - Kurla Road, Chakala, Andheri East, Mumbai, Maharashtra, 400093, India. If you would rather we came to you, say so — for a first session on content operations, your building is more useful than ours.


Questions from media and advertising clients

Do we legally have to label AI-generated content in India?

It depends who "we" are. The Information Technology (Intermediary Guidelines and Digital Media Ethics Code) Amendment Rules, 2026 - notified by MeitY on 10 February 2026 and in force from 20 February 2026 - place the labelling and metadata duties on intermediaries, meaning the platforms that host or generate the content, not on a production house or advertiser directly. Significant social media intermediaries must also ask uploaders to declare whether something is synthetically generated and verify that declaration automatically. So you will be asked to declare, and the platform labels on the basis of your answer. Build the ability to answer accurately for every asset you ship. Your counsel sets the requirement; we make the pipeline meet it.

What does the ASCI draft actually require?

It sorts synthetic content into three tiers rather than applying one rule to everything. High risk - misleading or infringing material - stays prohibited whether or not it carries a label. Medium risk, where the synthetic element could materially influence a consumer decision, expects disclosure using wording like "Audio/Video created using AI". Low risk, including colour correction, noise reduction and decorative backgrounds, needs none. Source: ASCI, Draft Guidelines for Responsible Labelling of Synthetically Generated Content in Advertising, 8 May 2026, still draft and self-regulatory.

Is the ten per cent watermark rule real?

Not in the rules as notified. The October 2025 draft did propose that a visible label cover a defined share of the display area and an audio disclosure a defined share of the duration. Those numeric thresholds were dropped before notification and replaced with a qualitative standard: prominent, easily noticeable and adequately perceivable. If someone is quoting you a percentage, they are quoting a superseded draft. What did survive matters more for a pipeline - permanent metadata or a unique identifier, embedded and not removable or alterable by users.

Can you dub our back catalogue into Indian languages?

Usually, though the honest scoping question is which parts of it are worth dubbing rather than how many hours we can process. Long-tail library content justifies an automated pipeline with light review. Flagship drama does not, and we will say so. See our AI video localisation page for how the pipeline and its review gates are structured.

Do you replace our editors and copywriters?

No, and projects starting from that premise are the ones that fail. What automates well is the repetitive, low-judgement layer: resizing, versioning, conforming, transcription, metadata, routing an asset to the right approver. The judgement layer is where your margin is. A team whose editors spend their week on exports rather than edits is the argument for automation, not against it.

Can you work inside the tools we already use?

That is normally the requirement rather than a bonus. You have an asset manager, an edit suite, a review tool and a scheduler, and the value is usually in the joins between them rather than in replacing any of them. We build against whatever APIs and webhooks exist, and where a system has none we say so in week one rather than week six.

Who owns the pipelines, prompts and models you build?

You do, on delivery: pipeline code, prompts, evaluation sets and configuration. We retain no licence to reuse your material and we do not subcontract, so no third party holds a copy of your assets you did not agree to. Where a build depends on a third-party model or API, we name it in the proposal along with what its terms say about your inputs.

What does this cost and how long does it take?

A scoping diagnostic is a short fixed piece of work that maps where the hours actually go and recommends what to automate first. Build work then runs as fixed-scope phases with defined exit criteria, not an open-ended retainer. Content operations phases are usually the shortest and pay back fastest. Localisation and provenance work takes longer, because the review stage is real engineering. The scoping call is free, and a fair share end with us naming an off-the-shelf tool instead.

Bring us the part of the week nobody wants to do

The best first session is not a capability pitch. It is two hours watching how an asset actually gets from edit to publish, and a straight answer on which parts of that are worth automating.

Start a conversation See our work