How to run an internal AI hackathon: a playbook for business teams

    An internal AI hackathon takes a small number of real business problems, puts cross-functional teams on them with tools like Microsoft 365 Copilot, ChatGPT, Claude or Gemini, and gives them a fixed window (one day to a week) to build something that works. Done well, it compresses weeks of "should we try AI for this?" into a few days of proof. It will not produce a production system, but it will tell you, quickly and cheaply, which problems AI can actually solve for your business and what the solution looks like.

    Can an AI hackathon really solve business problems quickly?

    Yes, for the right kind of problem, and at prototype speed. A hackathon is a way of finding out what works before you commit budget, not a way of skipping the commitment.

    The problems that suit a hackathon share three features. They are well defined: a specific team, a specific task, a specific output. The data or documents needed are available and safe to use. And a working demo can be judged in ten minutes by someone who does the job. Examples: drafting first-response emails to a category of customer query, turning weekly ops reports into a management summary, or checking supplier invoices against purchase orders.

    Problems that do not suit a hackathon: the vague ones ("use AI to improve customer experience"), the ones needing data nobody can release, and the ones where the hard part is integration with a core system. Put those through a proper pilot with engineering support.

    What you get at the end is a set of prototypes, a clear view of which deserve investment, and a team that now knows how to build with these tools. The gov.uk employer guide "Skills for AI: what works" notes that people learn AI skills fastest when applied to their own work; a hackathon is the most concentrated form of that.

    How do I run an AI hackathon for my company? Eight steps

    These are in order. Skipping the first three is the most common reason hackathons produce nothing usable.

    Step 1: Pick the problems

    Start four to six weeks out. Ask team leads for problems that are costing time, not for ideas about AI. Collect twenty or thirty, then cut to five to eight using the test above: specific, data available, judgeable by the person who does the job.

    Write each up as a problem statement (see below), with a business owner who will be available during the event and sit on the judging panel.

    Step 2: Form the teams

    Teams of three to five, and mixed. Each team needs someone who does the job the problem describes, someone comfortable with the tools, and ideally someone from a neighbouring function who will ask the obvious questions. The person who knows the process is worth more than the person who knows the prompt.

    Ten to forty participants is the workable range. Above forty, run two events or split into tracks with separate judging.

    Step 3: Set the guardrails

    Decide, in writing, before the event: which data can be used, which tools are approved, where outputs are stored, and who signs off before anything built is used in live work. Get it reviewed by information security and data protection, and circulate it a week ahead. The detail is in the guardrails section below.

    Step 4: Choose the tools

    Use what your company already licenses. A Microsoft 365 shop builds in Copilot, Copilot Studio and Power Automate; teams with ChatGPT Enterprise, Claude for Work or Gemini for Workspace use those. Every tool in play must have an enterprise agreement that keeps your data out of training sets. Consumer accounts are out.

    Provide a short list of approved building blocks: the chat tool, a way to build a simple agent, a way to connect to documents, and a low-code automation tool. That covers almost every business problem.

    Step 5: Timebox it

    Set the format (one, three or seven days; see the table below) and fix the schedule to the hour. Teams work better with hard checkpoints: problem restated by 10:00, first rough version by lunch, second by mid-afternoon, demo rehearsal at 16:00. Long unstructured stretches produce polished slides and unfinished builds.

    Step 6: Coach

    Every team needs someone who can get them unstuck in five minutes, usually a mix of internal champions and an external facilitator. Coaches walk the floor, not the stage. Their job is to stop teams building the wrong thing well: test with real examples early, strip features, get to a demo a sceptic can watch.

    Step 7: Demo and judge

    Demos are seven minutes each, with a live run on an example the judges supply, not one the team rehearsed. Judges score against published criteria (table below) and the business owner of each problem sits on the panel. Announce the criteria at the start of the event so teams build for them.

    Keep the prize modest. A big cash prize gets secrecy; a promise that the winning prototype will be taken forward gets collaboration.

    Step 8: Package what works

    Within a week, each prototype worth keeping is written up in the same short format: the problem, the build (prompts, agent configuration, workflow steps), test results, guardrails, and what it would take to run properly. This is the difference between a hackathon and a nice day out. Hand the packages to budget holders with a recommendation on each: adopt, pilot, park.

    What does the timeline look like for each format?

    AI hackathon timelines
    One-day hackThree-day buildDay Seven Sprint (seven days)
    Before the eventProblems chosen and written up; teams formed; guardrails circulated; tool access checkedSame, plus a one-hour briefing call per team the week beforeSame, plus a half-day problem-framing session with business owners two weeks ahead
    Day 109:00 brief and criteria; 09:30 teams restate problems; build through the day with two checkpoints; 16:00 demos; 17:00 judging and next stepsBrief, criteria, team kick-off; teams map the current process and agree what "working" means; first rough build by end of dayBrief, criteria, kick-off; teams shadow the actual job for the morning, collect real examples, map the process; afternoon: first prompts and tests
    Day 2Build day. Coaches push for tests on real examples by lunch; mid-afternoon peer review across teamsBuild day. First end-to-end version by close, however rough
    Day 3Morning: harden and rehearse; afternoon demos; judging; adopt/pilot/park decisionsTest with the people who do the job. Collect failures. Rebuild around them
    Day 4Build day. Connect to real documents or data sources under the agreed guardrails
    Day 5Second round of user testing; measure time or error rate against the current way of working; write up results
    Day 6Harden, document, prepare handover pack; rehearse demos
    Day 7Demo day: live runs, judging, decisions, and a written package per build handed to sponsors
    Week afterWrite-ups; adopt/pilot/park decisionsOwners confirmed for anything adoptedPilots scheduled for anything adopted; guardrails signed off
    Best forBuilding confidence and finding candidate problemsProducing two or three prototypes worth pilotingTaking one or two real problems to a working, tested build

    One day is about learning and discovery. Three days produces prototypes. The seven-day Sprint produces something a team can use the following Monday, subject to sign-off.

    How should we judge the entries?

    Publish the criteria before the event and score each on a 1 to 5 scale. Include the business owner of each problem on the panel, and at least one judge with no stake in any entry.

    AI hackathon judging criteria
    CriterionWhat judges are looking forSuggested weight
    Does it work?A live run on an unseen example gives a usable result30%
    Does it solve the stated problem?The build addresses the problem statement, not an easier neighbouring one20%
    Would the people who do the job use it?Fits how the work is actually done; the person who does the job today says they would use it20%
    Is it safe to use?Operates within the data and tool guardrails; failure modes identified; a human checks the output where it matters15%
    Could it be reused?The pattern (prompt, agent, workflow) could be applied to similar problems elsewhere10%
    Is it clearly explained?A seven-minute demo that a sceptical manager could follow5%

    Note what is missing: "innovation", "wow factor", "use of the newest model". Those criteria reward spectacle over usefulness.

    What makes a good problem statement?

    A good problem statement fits on half a page and answers five questions: who does this today, what do they do, how long does it take or how often does it go wrong, what would "better" look like in measurable terms, and what data or documents are involved. It names an owner. It does not mention a solution.

    Three example statements, written for ordinary business functions. These are illustrations, not client work.

    Operations. "The planning team receives around 40 delivery exception reports a day from depots as free-text emails. A coordinator reads each one, classifies it (late, damaged, address issue, other), and logs it in the tracker. This takes about two hours a day. Better would be a first-pass classification and a draft tracker entry that the coordinator reviews and corrects. Data: the emails (no customer personal data beyond an address), the tracker template. Owner: planning manager."

    Finance. "Accounts payable receives supplier invoices as PDFs and checks each against the purchase order: supplier, amounts, line items, PO number. About one in eight has a mismatch that needs a query back to the supplier. Checking takes roughly four minutes per invoice. Better would be an automated first check that flags mismatches and drafts the supplier query. Data: sample invoices with supplier details, PO extracts. Owner: AP team lead."

    Customer service. "The support inbox gets a steady flow of questions about order status, returns policy and delivery times. Agents answer from a knowledge base but each reply is written from scratch. Better would be a draft reply, grounded in the current knowledge base, that the agent edits and sends. Measure: time to first response and whether agents accept or rewrite the draft. Data: the knowledge base and anonymised sample queries. Owner: customer service manager."

    Each can be built in a day with a chat tool, a document connection and a few sample records. Each has an owner and a measure.

    What guardrails do we need?

    Guardrails are what let you say yes to using real examples, which is what makes the results credible.

    Data. Classify what teams can use: public information, internal but non-sensitive documents, and personal data. The first two are usually fine within enterprise tools. Personal data needs a decision from your data protection lead and, typically, anonymised or synthetic samples for the event. The ICO's guidance on AI and data protection is the reference point in the UK. Nothing commercially sensitive leaves approved tools.

    Licences and tools. Only tools on enterprise agreements. Confirm in writing with IT that the licences in use do not train on your inputs and that outputs are retained in your tenant. Consumer ChatGPT, personal Gemini accounts and free tiers of anything are excluded.

    Access. Teams get access to the documents and systems their problem needs and nothing else, set up the week before.

    Sign-off before anything goes live. Nothing built at a hackathon touches live customers, live finance or live HR processes without a named sign-off. Decide in advance who that is: usually the business owner plus whoever owns information security, and for anything involving personal data, the data protection lead. State it on the first slide of the event so nobody is surprised.

    Human in the loop. Every prototype should show where a person checks the output. In the three examples above, the coordinator, the AP clerk and the agent all review before anything is sent or logged. For most business uses that is the right design, not a temporary measure.

    If your organisation has, or is working toward, an AI management system such as ISO/IEC 42001, the hackathon guardrails should be drawn from it rather than invented for the day. The NIST AI Risk Management Framework is a useful checklist for the categories of risk to consider.

    What goes wrong at AI hackathons?

    Problems chosen for the AI, not the business. Teams pick something that shows off the tools rather than something anyone needs. Fix: problems come from team leads, with an owner and a measure.

    Slides instead of builds. The team spends the afternoon on a deck. Fix: live demos on judge-supplied examples, no slides longer than three pages.

    No one who does the job on the team. The build is elegant and nobody would use it. Fix: team composition rule in Step 2.

    Data panic on the morning. Nobody agreed what could be used, so teams use nothing real and the demos prove nothing. Fix: guardrails circulated a week ahead.

    Enthusiast capture. The three keenest people do all the building and everyone else watches. Fix: coaches rotate roles and check the least confident person has driven the tool.

    The winning build is never seen again. The most common failure of all. Fix: Step 8, with a named owner and a date for the adopt/pilot/park decision.

    What should we do the week after?

    The week after the hackathon matters more than the event. Four things, in this order.

    Write up every build in the standard package, including the ones that did not work; a documented dead end saves the next team a day. Make the adopt/pilot/park decision on each with the business owner and budget holder in the room. For anything adopted, confirm the guardrails, the sign-off and a named owner. For anything piloted, set a measure and a review date.

    Then tell the company. A short internal note with the three best builds and what happens next surfaces the next round of problems, because people write in with "we have one like that".

    How does Day Seven facilitate an AI hackathon?

    Day Seven runs internal AI hackathons for UK business teams in three formats: a one-day hack, a three-day build and the seven-day Day Seven Sprint, which takes a real internal problem to a working, tested build.

    Before the event we run the problem-framing work with your team leads, help write the problem statements, and agree the guardrails with your IT and data protection people. On the day, our trainers coach teams in whichever tools you have licensed (Microsoft 365 Copilot, ChatGPT, Claude, Google Gemini) and in the workflow and automation layer around them. We supply the demo-day structure, the judging criteria and the panel process. Afterwards we package each build so it can be reused or handed to engineering.

    Hackathons sit in the Prove stage of the Day Seven method: after leadership has set direction (Align) and teams have had hands-on training (Enable), a sprint is how the organisation gets something measurable to show for it before scaling (Embed). We run them on-site anywhere in the UK or live online. Details are on the AI hackathons page.

    Frequently asked questions

    How long should an internal AI hackathon be?

    One day to build confidence and find candidate problems. Three days for prototypes worth piloting. Seven days to take one or two problems to a working, tested build.

    How many people should take part in an AI hackathon?

    Between ten and forty, in teams of three to five. Below ten you have too few teams for the demos to be interesting. Above forty, split into tracks with separate judging or run two events.

    Do participants need to be technical?

    No. The most useful team member is the person who does the job the problem describes. Enterprise AI tools are built for non-technical users, and coaches handle the moments when someone gets stuck.

    Which AI tools should we use for a hackathon?

    The ones your company already licenses on enterprise terms: Microsoft 365 Copilot, ChatGPT Enterprise, Claude for Work or Gemini for Workspace, plus a low-code automation tool. No consumer accounts; the data terms are different.

    Can we use real company data in a hackathon?

    Internal, non-sensitive documents usually yes, within enterprise tools. Personal data needs a decision from your data protection lead and is normally replaced by anonymised or synthetic samples. Decide this in writing a week before the event.

    What should the prize be?

    Something small, plus a commitment that the winning build will actually be taken forward. Large cash prizes make teams secretive; a promise of follow-through makes them collaborate.

    Who should judge an AI hackathon?

    The business owner of each problem, a senior sponsor, someone from IT or information security, and at least one judge with no stake in any entry. Publish the criteria beforehand.

    What happens to the prototypes afterwards?

    Each one gets a short written package (problem, build, test results, guardrails, what it would take to run properly) and an adopt, pilot or park decision within a week. Adopted builds get a named owner and a sign-off before they touch live work.

    Can an AI hackathon replace AI training?

    No, but it is a strong follow-on. Teams get more from a hackathon if they have had hands-on training first, and the hackathon then shows what that training was for. See AI upskilling for how the two fit together.

    Should we run the hackathon ourselves or bring in a facilitator?

    Run it yourselves if you have confident internal champions, a clear problem list and someone with time to own the preparation. Bring in a facilitator if any of those is missing, if it is your first one, or if you want a tested build at the end. See hire an AI hackathon facilitator.

    Talk to us about your hackathon

    If you have problems in mind and want a team to take them to working builds, see the AI hackathons page or get in touch. Tell us what your team needs. We'll get back to you within one working day.

    Get in touch